← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

Principal, Data Scientist, Agentic AI Systems Engineering & Model Post-Training

Walmart · Bentonville, AR

📅 20/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

See the other 164,053 jobs in United States →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Position Summary... What you'll do...The Opportunity: Walmart’s Supply Chain AI Lab & Innovation Factory is building a new generation of production-grade agentic AI systems that reason over complex enterprise information, coordinate specialized agents, use tools safely, plan and execute long-horizon work, and continuously improve through rigorous evaluation and model post-training. This role exists to build those systems end to end—and to improve the models that power them. This is not an analytics-focused data science role. It is a deeply hands-on AI systems engineering position focused on designing, building, and operating production software. As a Principal Data Scientist in this space, you are a hands-on technical leader. You quickly turn hard, ambiguous problems into working full-stack prototypes—with a real user experience, APIs, telemetry, and an evaluation plan—then harden them into secure, reliable, observable, maintainable production systems. You independently own major product and platform domains within the shared agentic architecture, across the full technology stack—agent orchestration and model logic, backend services and APIs, the data layer, and web and CLI/TUI interfaces. What you will build and own: Build advanced agentic systems end to end: Design, build, test, launch, and operate production agentic applications and services with multi-step and long-horizon planning, tool use, retrieval, durable sessions and workflows, context management, multi-agent orchestration, human approval, and safe recovery when decisions or actions need intervention. Build a policy-first agent runtime and control plane with deterministic allow, deny, and ask decisions; least-privilege tool and data access; scoped identity and authorization; auditable human approvals; bounded subagent delegation; and safe cancellation, user steering, and retry behavior for long-running work. Build and own the full product path as a full-stack systems engineer: backend services and APIs, the data layer (database and schema design, data modeling, migrations, and streaming pipelines), telemetry, and accessible React/TypeScript experiences for associates and operators. Where the workflow demands it, design equally usable AI-native CLI/TUI or headless, structured-output interfaces that support automation, operations, and CI/CD-style integration. Where you will apply this: These capabilities will first be applied to some of Walmart's most complex operational and supply-chain problems, beginning with the Autonomous Supply Chain Engine and its Discovery Loop—a continuous, human-in-the-loop mechanism that monitors signals, discovers opportunities and risks, reasons through changing conditions, and generates strategic recommendations. Build and own key capabilities of the Discovery Loop within our Autonomous Supply Chain Engine. This is a flagship capability for this role: you will own major parts of its execution model, knowledge and context layer, evaluation loop, and learning path that turns feedback and business outcomes into measurable improvement. Prior supply chain or logistics experience is not required. You bring outstanding depth in agentic AI engineering, software engineering, and machine learning engineering, and you will work directly with AI, product, engineering, operations, data, and strategy partners on high-priority work. Engineer the agentic AI foundation: Own the design and build of advanced multi-agent harnesses, runtimes, and orchestration capabilities using frameworks such as Pydantic AI, LangGraph, LangChain, AutoGen, or LlamaIndex, or purpose-built custom infrastructure. Building your own runtime where that is the right call is a strength, not a gap. Establish typed contracts, testability, scalability, operability, and developer ergonomics as non-negotiable platform properties. Engineer explicit, inspectable agent execution. Design graph-based, state-machine-based, event-driven, planner/executor, or equivalent execution models as the problem warrants, with explicit state, typed dependencies and handoffs, conditional branching, fan-out and fan-in to subagents, retry and repair paths, loop detection, cost and time budgets, and hard termination conditions. Design external verification into the system—test runners, execution results, transaction outcomes, deterministic checks, and expert human review—because a model reviewing its own output is not verification. Build governed knowledge and context infrastructure for durable agent memory. Choose and combine the right representations—knowledge and context graphs, vector retrieval, relational and temporal models, or hybrids—with an explicit schema and ontology treated as a product contract. Own entity resolution and conflict handling, construction and enrichment pipelines from structured and unstructured sources, provenance and temporal validity, schema validation and evolution, and multi-hop retrieval that answers questions flat retrieval cannot. Build and operate agent skills, tool adapters, hooks and extension points, Model Context Protocol (MCP) clients and servers, structured outputs, function calling, enterprise APIs, identity, and secrets management. Own MCP lifecycle and reliability end to end: secure configuration and authentication, per-agent tool binding, schema compatibility, tool discovery, timeouts, health monitoring, retries, circuit breaking, quarantine, cleanup, failure isolation, and auditable operations. Create governed extension ecosystems for agents, skills, commands, plugins, hooks, tool adapters, and reusable workflows. Define stable contracts, compatibility and versioning strategies, secure installation and update paths, isolation boundaries, rollback behavior, and observability so extensibility does not become an uncontrolled code-execution surface. Treat agent quality as an engineering discipline: golden tasks, offline benchmarks, online experiments, adversarial and regression testing, failure analysis, measurable quality thresholds, and release gates. Establish practical AgentOps / LLMOps practices for prompt and tool versioning, tracing, evaluation datasets, workflow reliability, cost and latency controls, incident learning, and continuous improvement of long-running autonomous systems. Instrument the system so engineers and operators can answer, with evidence, what the agent did, why it was permitted, which model, tool, and policy version was involved, what data and integrations were used, what it cost, where it failed, and how to reproduce or remediate the outcome safely. Advance the models themselves Improve model reasoning and quality through hands-on post-training for the systems you own: reinforcement learning (RLHF/RLAIF), preference optimization, supervised fine-tuning, distillation, and reward and grader design to strengthen reasoning, tool-use reliability, and domain-specialized behavior across frontier models and smaller, domain-specialized models. Build the model-improvement flywheel. Turn production interaction traces, tool-use trajectories, human feedback, and successful and failed reasoning paths—together with Walmart's proprietary enterprise and operational data, synthetic data, and curated evaluation sets—into governed training and evaluation datasets. Use them to post-train, distill, evaluate, and redeploy increasingly capable domain-specialized models back into the agentic systems, so the platform you build continuously improves the models that power it. Own the data governance, provenance, privacy, and access controls that make this safe at enterprise scale. Design provider-aware, model-agnostic execution and routing layers that account for model capabilities—including multimodal inputs and outputs across text, images, and documents—context limits, structured outputs, streaming behavior, credentials, rate limits, transient failures, and explicit quality, latency, and cost trade-offs. Make every solution safe, scalable, and pr
Read the rest of the job →
For people searching Engineer

144,883 engineer jobs are open right now

Here's how to pick the right one and stand out in your application.

144.883Jobs
31.687IN
81%EN

That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.

Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.

Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.

When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.

👁 21 have read this
0 comments
Want to comment?

Leave your e-mail to comment, react and follow the posts for your role. It is free.

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →