← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

Head of Platform Reliability

Nava · Bangalore

📅 13/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

See the other 112,421 jobs in India →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
About Nava Nava is building next-generation AI infrastructure and inference platforms at global scale. We're looking for a Head of Platform Reliability to own the deployment, availability, and operational excellence of our AI infrastructure. This leader will be responsible for ensuring that our GPU clusters, platform services, and infrastructure operate reliably in production, while continuously improving automation, incident response, and platform resilience. What You'll Do Platform Reliability & Operations Own the day-to-day reliability, availability, and operational health of Nava's AI infrastructure platform.Lead production operations across GPU clusters, networking, storage, orchestration, and platform services.Define and drive Service Level Objectives (SLOs), Service Level Agreements (SLAs), and uptime targets.Ensure production environments meet the highest standards of reliability, scalability, and operational excellence. Deployment & Release Management Own the deployment strategy for software, firmware, infrastructure updates, and platform releases.Design and implement safe rollout mechanisms including phased deployments, canary releases, blue-green deployments, and rollback strategies.Ensure production changes are executed with minimal customer impact and operational risk.Establish deployment readiness reviews and operational change management processes. Incident Management Lead the incident response function for production systems.Establish incident management processes, escalation frameworks, and post-incident reviews.Drive Root Cause Analysis (RCA) and ensure corrective and preventive actions are implemented.Build a culture of operational learning and continuous improvement. Reliability Engineering Identify recurring operational issues and drive long-term engineering fixes rather than temporary workarounds.Improve system resilience through automation, observability, monitoring, alerting, and self-healing capabilities.Partner with engineering teams to eliminate reliability bottlenecks and technical debt.Drive capacity planning, performance optimization, and operational readiness. Cross-Functional Leadership Work closely with Platform Engineering, GPU Cluster Engineering, Networking, SRE, Product, and Customer Success teams.Ensure operational requirements are embedded into system design from the earliest stages.Drive operational excellence across multiple engineering functions. Team Leadership Build and lead a high-performing Platform Reliability and Site Reliability Engineering (SRE) organization.Mentor engineering managers and technical leaders.Foster a culture of accountability, ownership, operational discipline, and customer-first thinking. What We're Looking For 1015 years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Operations, or Cloud Operations.Proven experience leading large-scale production infrastructure teams.Deep expertise in operating highly available distributed systems and cloud platforms.Strong understanding of:Kubernetes and container orchestrationLinux systems administrationInfrastructure automationMonitoring and observability platformsIncident management and production operationsHigh-availability architecture and disaster recoveryExperience managing large-scale production deployments and release management.Strong knowledge of operational metrics including SLAs, SLOs, SLIs, MTTR, and incident response best practices.Exceptional leadership, stakeholder management, and communication skills. Nice to Have Experience operating AI infrastructure, GPU clusters, or large-scale inference platforms.Familiarity with NVIDIA GPU infrastructure, InfiniBand/RDMA networking, and distributed AI workloads.Experience with GitOps, Infrastructure as Code (Terraform, Ansible), CI/CD pipelines, and production automation.Exposure to cloud-native platforms, HPC environments, or hyperscale infrastructure. Why Join Nava Shape the future of AI infrastructure - Lead reliability and operational excellence for one of the worlds most advanced AI platforms.Scale with purpose - Build and grow production systems that power next-generation AI workloads for global customers.Collaborate with world-class talent - Work alongside top-tier engineers, researchers, and operators solving some of the hardest infrastructure challenges in AI.Build and lead at the frontier - Establish the operational foundations of Navas global AI platform while growing a high-impact reliability engineering organization. Skills: drive,platforms,infrastructure,availability,reliability,operations,reliability engineering,operational excellence,automation,management .
Read the rest of the job →

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →