← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

Research Engineer Agent Architectures (Coding & Autonomous Systems)

AryaXAI · Mumbai City

🌐 Remote📅 11/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

See the other 113,540 jobs in India →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Lexsi Labs is the leading frontier AI lab focused on building aligned, interpretable, and safe superintelligent systems. While that is the vision, the mission to build safety aware autonomous system in the extreme near term. Our research work spans around areas like AI alignment methodologies, interpretability led system design, and foundational model research across structured, tabular, and new autonomous system designs. We published about 25+ papers in the past 15 months across leading conferences ICLR, ICML, WWW, IJCNN, MICCAI, Eurips etc. Our labs are located in India (Mumbai & remote), Paris, & London. We operate with a flat structure, high autonomy, and a strong bias toward engineers who take full ownership of what they build, from architecture to production behavior. The Role Most of the variance in agent performance comes from the system around the model, not the model itself. This role researches that system. The coding agent is the primary testbed: it takes an objective, works on a repository, verifies its own changes, and produces a record of what it did. It runs inside a customer's network, usually with no internet egress, against codebases that are large, old, thinly tested and load-bearing. Small models under real latency and cost budgets are a target, not a fallback. Three areas: Harness. Action space and tool surface design, context construction policy, where deterministic program analysis should replace model inference, control topology (single loop versus decomposition), verification design, and allocation of inference-time compute. Memory. Schemas for execution state held outside the context window, compaction policy and what it destroys, retrieval under closed-world constraints, and whether an agent measurably improves on a repository over time. Evals. Task construction from real repository history with executable verification, scoring partially-checkable long-horizon work, variance and contamination control, and failure taxonomies that attribute a failure to a component. The work is controlled experiments on architecture, not prompt tuning: isolate a failure from eval traces, form a hypothesis about the responsible component, change it, ablate it, and establish whether the gain holds across repositories and model sizes. Outputs are papers, benchmarks and released tooling where we can publish, and shipped architecture where we cannot. You will work directly with our alignment and interpretability researchers on post-training, behavioral evaluation, and reading what an agent actually did rather than what its trace claims. What We Are Looking For Research judgment with systems depth. You should be able to isolate a failure, design the experiment that tests your explanation of it, build what the experiment needs, and tell the difference between a real gain and an artifact of your setup. Agentic systems. Two or more years building agents that ran against real workloads, not demos. Working knowledge of the current landscape (ReAct-style agents, LangGraph, LangChain, Semantic Kernel, the current generation of coding agents) and a specific account of where each stops working. Experience with tool-use protocols and orchestration under partial failure, retries, timeouts and non-idempotent actions. Evaluation. You have built an eval dataset or harness that other people then used. Comfortable with SWE-bench-class benchmarks and their construction, containerised task execution at scale, statistical treatment of noisy multi-run results, and pass@k, best-of-N and majority-vote scoring and their failure cases. Code intelligence and program analysis. ASTs and tree-sitter, static and dataflow analysis, symbol indexing and code search, call graph and dependency resolution, codemods and automated migration tooling. You should be able to build a repository-scale index and defend its design. Backend and infrastructure. Advanced Python. Sandboxing and containerisation (Docker, gVisor, Firecracker or equivalent), distributed execution of thousands of parallel trials, artifact and dependency management, and on-premise or air-gapped deployment. Experiments that cannot be run at volume are not useful here. Observability. Distributed tracing, structured logging, OpenTelemetry, and the replay and inspection tooling that makes a non-deterministic system debuggable end to end. Model side. You do not need to be a training specialist, but you should be fluent in post-training methods (SFT, DPO, RL for agents), distillation, inference-time scaling, and quantisation and serving trade-offs at small parameter counts. You should have a view on where architecture ends and the model begins. You treat performance, reliability, cost, safety and interpretability as one connected set of constraints, and you make reasonable calls when the problem is loosely specified and ownership is assumed rather than assigned. Qualifications Five or more years in research engineering, systems, or ML .
Read the rest of the job →
For people searching Engineer

144,883 engineer jobs are open right now

Here's how to pick the right one and stand out in your application.

144.883Jobs
31.687IN
81%EN

That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.

Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.

Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.

When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.

👁 21 have read this
0 comments
Want to comment?

Leave your e-mail to comment, react and follow the posts for your role. It is free.

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →