🎁 Antes de postularte, practica esta entrevista. Crea tu cuenta gratis en WorkMundi y recibe un Entrenamiento de Entrevista en HelpsYouSpeak — sin costo, sin tarjeta. Quiero mi entrenamiento →
Senior Platform Reliability & AI Operations Engineer Most reliability teams react to incidents. We engineer them out of existence. We operate an AI-native community and social engagement platform that underpins customer relationships for Fortune 100 brands. When our platform goes down, it isn't a blip on an internal dashboard — it's a reputational crisis for some of the most recognized companies on earth. That context shapes everything about how we work, what we build, and who we hire. We're looking for a senior reliability engineer who can carry production on their shoulders and build the autonomous systems that progressively carry it for them. You'll join a team where AI agents are first-class operational teammates — triaging alerts, validating changes, drafting root-cause analyses, and applying remediations within defined guardrails. Your mission is to make that surface area grow every single week. Your Day-to-Day Carry the pager and own the outcome. You're the first responder on your shift window. When production degrades, you command the incident — diagnose, mitigate, escalate when blast radius demands it, and restore service. You treat every customer-impacting minute as personal accountability. Engineer autonomous operational workflows. Build, deploy, and refine the AI agents that handle pre-triage, change-gate validation, auto-healing, RCA drafting, and preventive-fix tracking. The agents are the product; your operational expertise is the training data. Ship safe production changes. Every deploy, config update, and cost-optimization action flows through quality gates with a validated rollback plan. You abort without hesitation the moment telemetry deviates from the expected path. Investigate to true root cause — then close the loop. Separate symptom from cause with disciplined analysis, then go further: identify the systemic prevention, build it, and track it to production. Unshipped RCA action items are unfinished work. Generalize every manual intervention. A one-off fix restores service; encoding it into an agent, runbook, or guardrail prevents recurrence. You're measured by how much the autonomous layer can handle — not by how many tickets pass through your hands. Multiply team knowledge. Encode procedures, context, and decision logic so agents can retrieve it and the next responder never starts from scratch. In a distributed, async organization, undocumented expertise doesn't count. Who You Are 5+ years of hands-on production operations in SRE, Platform Engineering, DevOps, or Cloud Infrastructure at SaaS scale — with real first-responder incident experience, not adjacent project work. Battle-tested on AWS — multi-AZ, multi-account environments, infrastructure-as-code, production incident management, and change control with gates and rollbacks. You've weathered significant outages and carry the operational intuition that only comes from living through them. Radically self-directed. You run your shift like a founder runs a company. You identify gaps, prioritize ruthlessly, ship solutions, and raise the bar — without waiting for direction. If a standard is wrong, you challenge it openly; you never quietly ignore it. AI-native in practice, not in theory. You routinely delegate substantive operational work to agents, critically evaluate their output, and iterate on the underlying capabilities when they fall short. Experienced with agentic tooling — Claude Code, Codex, Warp, custom agent frameworks — and compulsively curious about new models and techniques as they emerge. AWS Solutions Architect – Associate or higher — or a production track record that renders the certification a formality. Fluent, precise English — in incident-bridge communication and in long-form writing alike. Committed to shift-based coverage. On-call rotations and your designated shift window are foundational to the role, with the time-zone overlap your window requires. OFAC-clear country of residence. Bonus Points Original contributions to the agentic operations or AIOps space — open-source tooling, technical writing, conference talks, or shipped internal platforms. Production experience with multi-tenant B2B SaaS — community platforms, social tools, customer-experience products, or observability systems. Working knowledge of Grafana, Prometheus, Datadog, PagerDuty, or OpsGenie. Azure exposure alongside your AWS depth. Evidence of deep, sustained obsession with a hard problem — professional or personal. Depth of curiosity matters more than breadth of résumé. What You'll Walk Away With You won't just read about the future of AI-driven reliability — you'll be one of the engineers who built it, on a platform that Fortune 100 brands depend on daily. The agent-ops skills, incident patterns, and architectural instincts you develop here are things the broader industry is still struggling to define, let alone hire for. How We Work Enterprise clients, startup velocity. Fortune 100 contractual stakes with a small-team cadence — weekly delivery cycles, fast decision-making, and a playbook that evolves as the field does. Uncapped tooling and compute. The agent harness is the product. If the right answer is a bigger model, more infrastructure, or a tool we haven't adopted yet, we invest. Fully remote. Global team.
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.