🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
We are looking for an LLM / Agentic Evaluation Rig Engineer to build the system that decides whether our AI output is good enough to ship. Because our commentary sits next to externally reported financials, we cannot rely on vibes grounding, faithfulness, and hallucination have to be measured, tracked, and gated before anything reaches a customer. You own the evaluation infrastructure: the datasets, the scorers, the harnesses, and the CI gates that hold the AI and agentic layers to a hard quality bar. You are the team's source of truth on whether a model, prompt, or agent change is actually an improvement and the one who blocks it if it isn't. What makes this role different You define "good enough to ship" your gates block regressions in grounding and faithfulness from reaching production. Evidence over vibes every claim is checked against the verified source data it must be grounded in. Agentic evaluation you evaluate multi-step reasoning flows, not just single prompts. Real leverage your rig is how the whole AI team moves fast without breaking trust. Responsibilities Datasets & Scorers (35%) Build and curate evaluation datasets, including adversarial and edge-case sets with ground-truth labels Build scorers for grounding, faithfulness, hallucination, factual consistency, and structured-output validity Combine rule-based checks, reference-based metrics, and LLM-as-judge where appropriate Verify generated claims map to verified source data no unsupported statements Harnesses & CI Gates (30%) Build harnesses that run evaluations reproducibly across model, prompt, and agent versions Wire evaluation into CI so grounding / faithfulness regressions block releases Track quality over time with dashboards and clear pass / fail thresholds Agentic Evaluation (25%) Evaluate multi-step / agentic flows routing, tool-use, verification, confirmation Build trace capture and step-level scoring for agent runs Detect where a flow silently degrades Collaboration (10%) Partner with the Staff AI Engineer to turn findings into model / prompt / orchestration improvements Partner with QA to integrate AI evaluation into the broader release process Technical Stack Evaluation LLM eval frameworks (promptfoo, DeepEval, Ragas, LangSmith) LLM-as-judge, reference-based metrics Dataset / ground-truth curation AI & Orchestration LLM APIs & managed LLMs (Bedrock / Vertex / Azure OpenAI) RAG & agentic patterns (LangGraph) Structured-output validation Engineering Python CI/CD (GitHub Actions) Dashboards & metrics tracking What You'll Build in Year One A labeled evaluation dataset suite (including adversarial cases) for the generation and agentic layers. A scorer library for grounding, faithfulness, hallucination, and structured-output validity. A reproducible harness wired into CI that blocks releases on quality regressions. Step-level trace capture and scoring for agentic flows, with dashboards leadership can trust. Required Qualifications Core 4+ years in software / ML engineering, with hands-on work building LLM evaluation or quality tooling. Real understanding of grounding, faithfulness, and hallucination and how to measure them rigorously. Technical Strong Python and solid engineering practices (reproducibility, CI/CD). Comfort designing evaluation for non-deterministic systems without producing flaky or meaningless metrics. Familiarity with LLM eval frameworks and LLM-as-judge patterns. Nice-to-Have Experience evaluating agentic / multi-step LLM systems. Familiarity with RAG, structured output, and managed LLMs in-VPC. FinTech / financial-services domain or other high-stakes, correctness-critical AI. Background in statistics or measurement / metrics design. .
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.