🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Role: Platform SRE + AI Experience: 710 yrs Location: Bangalore Senior Platform SRE: Owns reliability and performance of the platform end to end leads incident response, drives cluster and runtime tuning standards, and sets direction for the AI/MCP platform.Expected to mentor and influence vendor and stakeholder decisions.Looking for a strong AWS & Kubernetes (EKS) expert to manage production-scale cloud infrastructure, platform reliability, performance tuning, and automation. Experience with AI/LLM platform operations, monitoring, troubleshooting, and cloud-native environments is highly preferred.This is a deep platform and performance-engineering role owning production Kubernetes, AWS infrastructure, and the reliability of LLM-backed and agentic services in a regulated environment.It goes well beyond standard SRE: we are looking for an engineer who tunes clusters, diagnoses memory and runtime behaviour, and operates AI platform workloads at production scale Mandatory Skills: Deep AWS expertise End-to-end production deployment on AWS is a must EC2, ECS, EKS, Lambda, S3, IAM, RDS, API Gateway, VPC, EFS, SNS, SQS, EventBridge, CodeBuild.Kubernetes (deep & mandatory) Production EKS ownership cluster configuration (CPU/ RAM sizing, node groups, autoscaling), workload fine-tuning (requests/limits, HPA/VPA, eviction policies), and hands-on Helm chart management.Memory expertise Container vs. runtime memory models, OOMKill diagnosis, cgroup behaviour, JVM/Python heap tuning, and database buffer-pool configuration.Serverless fine-tuning Lambda memory/concurrency/cold-start optimization, ECS/Fargate task sizing, and serverless cost-performance tradeoffs.Database operations Running and tuning RDS or equivalent in production; StatefulSet-based DB deployments in K8s a strong plus.AI platform & MCP environment Hands-on deployment or operation of LLM-backed services, MCP servers, or agentic pipelines on cloud infrastructure. Core Requirements: 510 years in software development, DevOps, SRE or production support roles.Working knowledge of GCP and/or Azure multi-cloud integrations, platform differences, hybrid workloads.CI/CD with GitHub Actions, Jenkins or GitLab.Scripting proficiency in Python and/or Bash; YAML fluency.Terraform and/or CloudFormation for infrastructure provisioning.ITIL framework across Incident, Change, Problem and CAPA management.Monitoring with Splunk and/or Grafana including infra-level resource and memory dashboards.ServiceNow and JIRA; strong ITSM discipline.Bachelors degree or equivalent practical experience .