← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

Senior Data Scientist - LLM Evaluation & Adversarial Data

Morpheus Talent Solutions · United States

📅 20/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

See the other 152,342 jobs in United States →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Public benchmarks saturate faster than they can be built, and many are contaminated by the time they are widely adopted. That leaves a gap: when a model scores well on everything, nobody can say where it actually breaks or what to train on next. We are hiring senior data scientists to close that gap . You will hunt for the failure modes of frontier LLMs, turn them into rigorous benchmarks that models cannot pass by memorization, and build the synthetic data pipelines that convert those findings into training signal. You will work fluently with current evaluation harnesses and know where their defaults mislead. This is adversarial, empirical work, and your success metric is how often you can make a strong model fail in a way that is real, reproducible, and diagnostic. What you will do 1. Failure mode discovery Systematically probe frontier models for weaknesses across reasoning, long context, tool use, code, and instruction following, as well as domain-specific tasks such as legal, clinical, financial, and engineering work. Move beyond anecdote: build taxonomies of failure, quantify how often each occurs, and separate genuine capability gaps from prompt formatting artifacts and evaluation harness bugs. Characterize where a failure originates, whether in pretraining coverage, post-training behavior, context length, decoding, or scaffolding, because the fix differs in each case. Run controlled experiments to show a failure is real and reproducible rather than incidental. 2. Hard benchmark construction Design and ship benchmarks where the strongest available models land in a useful accuracy band, hard enough to leave headroom but not so hard that scores are indistinguishable from noise. Own dataset quality end to end: item writing and sourcing, expert annotation and adjudication, inter-annotator agreement, gold-label verification, and honest error bars. Build contamination defenses in from the start, including private held-out splits, freshness-dated items, overlap checks against public corpora, and periodic re-verification as models refresh. Analyze results at the item level, so a benchmark tells you what a model cannot do rather than just what it scored. Ensure every item is verifiable through a programmatic checker, a closed-form answer, or a clear grading rubric. Unverifiable items do not ship. 3. Novel benchmark design Invent evaluation formats that do not exist yet, in areas where current benchmarks are weak or absent, including agentic multi-step workflows, tool use, long-horizon consistency, calibration, and cross-modal grounding. Write the design doc: what capability is being measured, why existing benchmarks do not measure it, and what a score does and does not tell you. Pilot fast and kill fast. Most benchmark ideas fail, and the skill is finding that out in two weeks rather than two quarters. 4. Synthetic data pipelines Build scalable generation pipelines using seed-task expansion, instruction evolution, self-play and debate, and program-synthesized items with verifiable ground truth. Design the filtering stack, which matters more than the generator: verifier models, execution-based checks, unit tests, deduplication, difficulty calibration, diversity sampling, and toxicity and PII screening. Manage the failure modes of synthetic data itself, including mode collapse, verifier gaming, and quality drift across pipeline versions. Instrument everything, from provenance and model versions to per-stage yield and cost per accepted item, so results are reproducible and cost is defensible. 5. Off-the-shelf dataset leverage Maintain deep working knowledge of the existing dataset landscape, covering reasoning and knowledge suites, code and agentic benchmarks, math, long context, multilingual, safety, retrieval, and domain corpora, including which ones are contaminated, mislabeled, or measuring something other than what their name implies. Adapt rather than rebuild where possible: perturb, recombine, harden, translate, extend to new modalities, or filter existing datasets into higher-difficulty subsets that recover discriminative power. Audit third-party datasets before anyone builds on them. Required 5+ years in ML or data science, with at least 2 years working directly on LLMs in evaluation, data curation, fine-tuning, or applied research. Demonstrated experience building datasets or benchmarks that others used, with a real quality bar covering annotation protocols, agreement statistics, adjudication, and versioning. Strong Python and the modern LLM stack: PyTorch, HuggingFace, an evaluation harness such as lm-eval-harness, HELM, Inspect, Harbor, or in-house, and comfort orchestrating large batched inference jobs across multiple model APIs. Solid experimental statistics: power analysis, effect sizes, paired designs, multiple-comparison discipline, and the judgment to know when N is too small to conclude anything. Adversarial instinct, meaning a track record of breaking systems that were assumed to work, and the rigor to distinguish a genuine break from a bad prompt. Clear technical writing. Strongly preferred PhD or MSc in a quantitative field. Published benchmarks, evaluation papers, or widely-used open datasets. Experience with LLM-as-judge methodology and its pathologies, including position bias, verbosity bias, self-preference, and how to validate a judge against human labels. Production synthetic data generation at scale, with measured downstream training impact. RLHF or RLAIF, preference data collection, or reward model evaluation. Deep expertise in a domain we can build hard evals for, such as mathematics, competitive programming, security, law, medicine, finance, or engineering. Experience managing annotation vendors or expert contractor pools.
Read the rest of the job →

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →