🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
At EY, youll have the chance to build a career as unique as you are, with the global scale, support, inclusive culture and technology to become the best version of you. And were counting on your unique voice and perspective to help EY become even better, too. Join us and build an exceptional experience for yourself, and a better working world for all. EY- Assurance Senior Digital Role: GenAI / Agentic AI Evaluation Engineer (Quality, Safety & Reliability) Position Details As part of EY GDS Assurance Digital, you will help design, build, and scale a standardized evaluation capability, focused on evaluating GenAI, RAG-based, and Agentic AI solutions before deployment. This role sits at the intersection of AI evaluation engineering, Responsible AI, and GenAI security/red teaming. The primary objective is to ensure GenAI/agentic systems are safe, reliable, robust, and fit-for-purpose, by designing evaluation strategies, building repeatable test harnesses, and generating auditable evidence that supports go/no-go decisions. You will work with global stakeholders (product teams, solution architects, risk & compliance, and assurance leadership) to define evaluation requirements, request test datasets from product teams, execute rigorous evaluations (functional + non-functional), and recommend mitigations and controls to reduce risk. This is a core full-time role that requires a hands-on AI Development mindset, strong evaluation mindset, and the ability to translate risk concerns into practical testing strategies and measurable acceptance criteria. Responsibilities Define and operationalize evaluation strategies for GenAI systems across use cases like Q&A assistants, summarization, extraction, drafting, agentic systems, and multi-step workflows. Translate business use-cases into a structured evaluation plan: scope, assumptions, success criteria, datasets, metrics, red-team scenarios, thresholds, and reporting requirements. Drive standardization: reusable evaluation templates, test case libraries, scoring rubrics, and reporting formats across product teams. Design structured dataset requirements for product teams and ensure coverage across: Core user journeys and primary business intents Edge cases (rare prompts, ambiguous queries, incomplete context) Adversarial cases (malicious prompts, jailbreak attempts, prompt injections) Bias & fairness cases (sensitive demographic proxies, protected attributes, stereotyping patterns) Define guidance for dataset sufficiency and statistical coverage (e.g., minimum samples, distribution balance, scenario matrices, stratification by intent/risk). Build reusable evaluation pipelines for: Answer quality (correctness, relevance, completeness, clarity) Grounding & faithfulness (RAG-specific: faithfulness, context precision/recall, hallucination rate, citation quality) Agentic behavior (tool-call accuracy, tool misuse, goal completion, step correctness, unnecessary actions, loop detection, safety of tool outputs) Operational quality (latency, cost/token budget, throughput, stability, retries, failure recovery) Combine LLM-as-judge and human evaluation in a calibrated way (rubric design, sampling plans, agreement checks). Implement automated evaluation harnesses in Python (preferred), enabling: batch runs on scenario suites configurable metric definitions reproducible runs with run IDs and artifacts storage of traces and outputs for auditability Execute structured red teaming aligned to OWASP Top 10 for LLM Applications, covering (examples): Prompt injection (direct + indirect) and tool hijacking Sensitive data disclosure / PII leakage Insecure output handling (downstream injection) Training data leakage / memorization probes Model denial-of-service / denial-of-wallet patterns Integrate evals into development lifecycle: pre-release regression gates, CI checks, benchmark comparisons across model versions/prompts/tools/retrievers. Perform adversarial testing for agentic workflows: tool misuse / over-permissioned tool access unauthorized action execution exfiltration via tools/connectors prompt injection via retrieved documents (RAG poisoning) Recommend mitigations: input validation, retrieval filtering, tool sandboxing, least-privilege permissions, guardrails, policy prompting, refusal logic, output encoding, monitoring alerts. Produce high-quality evaluation reports that are auditable and decision-ready, including: methodology, datasets, metrics, thresholds quantitative results qualitative results risk assessment summary and recommended control actions Present findings to stakeholders in a crisp, risk-informed manner; clearly explain residual risk, limitations, and rationale for go/no-go. Key Requirements/Skills & Qualification: Excellent academic background, including at a minimum a bachelors or a masters degree in data science, Statistics, Engineering, Operational Research, or other related field with strong focus on modern data .
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.