🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
About Dusker AI Dusker AI specializes in benchmarking and evaluating AI agents to help organizations understand real-world performance. Using expert-driven frameworks, we assess AI systems across reasoning, reliability, adaptability, and safety to ensure they are truly production-ready. From conversational AI to autonomous agentic systems, we build cutting-edge evaluation frameworks that enable organizations to develop trustworthy, high-performing AI solutions. Role Overview We are seeking an experienced RLHF / Post-Training Engineer to build the preference data pipelines, reward models, and alignment training loops that turn raw model capability into dependable behaviour. You will own supervised fine-tuning and preference optimization workflows end to end, from response sampling and reward modeling through training runs and post-training evaluation. You will work alongside evaluation scientists, domain experts, and infrastructure engineers to close the loop between measured model weaknesses and the training interventions that actually fix them. This role is ideal for someone passionate about preference learning, reward modeling, alignment methodology, and rigorous measurement of post-training gains. Key Responsibilities Design and run supervised fine-tuning and preference optimization pipelines using methods such as PPO, DPO, and GRPO. Build reward models and rubric-based graders that convert expert human judgment into reliable training signal. Architect preference data collection workflows covering response sampling, pairwise comparison design, and adjudication of disagreement. Develop reinforcement learning loops with verifiable rewards for reasoning, tool calling, and instruction following. Conduct systematic ablations to isolate which data mixtures, objectives, and hyperparameters genuinely drive gains. Collaborate with evaluation scientists to measure post-training effects on reasoning, safety, refusal behaviour, and regression risk. Optimize distributed training throughput and cost using LoRA, QLoRA, mixed precision, and efficient checkpointing. Guide annotation leads on rater calibration, quality control, and early detection of reward hacking in collected data. Document training recipes, data lineage, and experiment results so that every run is independently reproducible. Required Qualifications Bachelor's or Master's degree in Computer Science, Machine Learning, Statistics, or a related quantitative field. 4 or more years of machine learning engineering experience, including at least 2 years working directly with large language models. Strong Python skills and deep hands-on use of PyTorch alongside Hugging Face Transformers, TRL, PEFT, and DeepSpeed or Accelerate. Practical experience with post-training methods such as supervised fine-tuning, RLHF, DPO, or comparable preference optimization techniques. Demonstrated experience training models across multiple GPUs, including distributed strategies and memory optimization. Working knowledge of reward modeling and its failure modes, including reward hacking, over-optimization, and annotator bias. Familiarity with evaluation methodology, including held-out benchmarks, LLM-as-judge harnesses, and significance testing. Excellent written communication and the discipline to report results honestly, including the negative ones. Preferred Qualifications Experience with reinforcement learning from verifiable rewards or reasoning-focused post-training. Background in MLOps tooling such as Docker, Kubernetes, CI/CD, and experiment tracking with Weights and Biases or MLflow. Exposure to large-scale training on AWS, Azure, or GCP, including spot capacity and fault-tolerant checkpointing. Open-source contributions to training, alignment, or evaluation libraries. Publications in alignment, preference learning, or model evaluation. .
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.