← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

Senior Software Engineer - Model Training & AI Evals (Delhi)

Chegg India · Delhi

📅 24/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

See the other 113,540 jobs in India →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
ABOUT THE ROLE We are looking for a Senior Engineer to join our AI team at the intersection of evaluation science, post-training, and foundation model development. You will own our end-to-end eval and benchmarking infrastructure the critical feedback loop that drives every major model improvement while contributing hands-on to post-training pipelines for industry-specific vertical foundation models. This role is ideal for someone who has worked directly inside an LLM lab and understands what rigorous evaluation looks like at scale: designing the taxonomy of skills being measured, identifying failure modes, engineering synthetic data to close capability gaps, and translating eval signals into actionable training decisions. What You'll Do Evaluation & Benchmarking - Design and own task-level evaluation frameworks for LLM agents and base models, covering multi-step reasoning, tool/API use, instruction following, and domain knowledge grounded in real user failure modes rather than off-the-shelf benchmark suites. - Build comparative benchmarking pipelines to assess leading frontier models (GPT-4o, Gemini, Claude, Llama, Mistral, etc.) against each other and against internal models, with structured analysis of where each model family fails, regresses, or excels across subjects, topics, and task types. - Produce capability gap reports that quantify performance deltas across dimensions such as subject-matter accuracy, reasoning depth, factual consistency, and refusal behaviour. - Track model version regressions across provider releases to maintain a living competitive intelligence benchmark. - Develop domain-specific benchmarks tailored to vertical use-cases (e.g., STEM tutoring, legal, finance, healthcare) including problem taxonomy design, rubric definition, and inter-annotator agreement pipelines. - Define and drive synthetic data generation strategies to systematically address model shortcomings in specific subjects, topics, and skill areas: - Identify low-performance clusters from eval results and translate them into targeted data generation prompts and pipelines. - Design LLM-assisted pipelines for generating high-quality, diverse, and verifiable synthetic training and evaluation data at scale. - Validate synthetic data quality through auto-eval, human review, and downstream model performance lift experiments. - Build automated regression suites integrated into CI/CD workflows to detect capability degradation across fine-tuning runs and model updates. - Partner with product, curriculum, and research teams to translate eval insights into prioritized post-training and data flywheel decisions. Post-Training & Fine-Tuning - Lead or directly contribute to SFT, RLHF, RLAIF, and DPO training runs on industry-specific vertical foundation models from dataset design through training execution and eval-gated release. - Curate and engineer high-quality instruction-tuning and preference datasets for domain adaptation, with hands-on experience distinguishing signal from noise in annotation pipelines. - Define data quality criteria, rejection sampling strategies, and deduplication pipelines for SFT corpora. - Design preference pair construction methodologies and reward model training setups grounded in domain-specific quality rubrics. - Implement and experiment with alignment techniques including reward modelling, process reward models (PRMs), and constitutional/RLAIF approaches. - Run ablation studies and controlled experiments to attribute model behaviour changes to specific data or training interventions not just report final numbers. - Contribute to continual pre-training and domain-adaptive fine-tuning pipelines for vertical models, including domain data sourcing, mixing strategies, and curriculum design. Infrastructure & Tooling - Build scalable eval pipelines that run automatically on every training checkpoint and integrate into CI/CD for continuous model quality tracking. - Maintain model cards, eval leaderboards, and internal dashboards providing visibility across experiments for both technical and non-technical stakeholders. - Ensure reproducibility through rigorous experiment tracking (W&B;, MLflow, or equivalent), versioned datasets, and documented training configs. Required WHO YOU ARE - 5+ years of ML/AI engineering experience, with at least 23 years focused on large language models. - Lab pedigree: Direct, hands-on experience at an LLM lab, AI research organization, or equivalent frontier AI team you have shipped models, not just called APIs. - Familiarity with the full model lifecycle: pre-training data, post-training alignment, eval, and production deployment. - Deep practical expertise in post-training methods: - SFT, RLHF, RLAIF, DPO, PPO from dataset construction through training and eval-gated release. - Experience with reward modeling, preference data curation, and quality control for alignment pipelines. - Demonstrated experience designing LLM evaluation frameworks beyond .
Read the rest of the job →
For people searching Engineer

144,883 engineer jobs are open right now

Here's how to pick the right one and stand out in your application.

144.883Jobs
31.687IN
81%EN

That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.

Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.

Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.

When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.

👁 21 have read this
0 comments
Want to comment?

Leave your e-mail to comment, react and follow the posts for your role. It is free.

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →