🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Role Summary We are seeking experienced AI-OPS Engineers with a strong blend of Site Reliability Engineering (SRE), Cloud Operations, Automation, Observability, and Generative AI expertise. Candidates should be capable of designing and implementing intelligent operational solutions that improve incident detection, diagnosis, remediation, and operational efficiency. Required Technical Skills AI/ML & Generative AI OpenAI / GenAI solutions on GCP or AWS Machine Learning fundamentals MLOps AI Agents / Agentic AI Retrieval Augmented Generation (RAG) Enterprise Search AI Governance Prompt Engineering Programming & Automation Python (strong requirement) PowerShell Ansible Terraform Infrastructure as Code (IaC) DevOps practices Cloud & Platform Engineering GCP and/or AWS Kubernetes Container-based platforms Cloud-native operational tooling Observability & AIOps Splunk Observability platforms and monitoring tools Incident Analytics Event Correlation RCA Analytics Predictive Alerting Self-Healing Automation SRE & IT Operations Reliability Engineering Incident Management Production Support ITSM platforms (Remedy preferred) Problem Management Operational Excellence Preferred Experience 7+ years in Infrastructure Operations, SRE, Platform Engineering, DevOps, or AIOps 3+ years working with Cloud Platforms (AWS/GCP) Experience implementing automation and self-healing solutions Experience building AI-assisted operational workflows Experience with enterprise monitoring and observability platforms Preferred Certifications AWS Certified Solutions Architect / DevOps Engineer Google Professional Cloud Architect Certified Kubernetes Administrator (CKA) ITIL Foundation AI/ML or Generative AI certifications Key Responsibilities Design and implement self-healing operational workflows. Develop AI-assisted RCA and operational intelligence capabilities. Build and maintain knowledge and runbook copilots. Improve monitoring, observability, and incident response processes. Automate operational tasks using Infrastructure as Code and orchestration tools. Collaborate with SRE, platform, and application teams to improve reliability and operational efficiency. Evaluate and implement Agentic AI solutions for autonomous operations. Ideal Candidate Profile Candidates should possess a strong combination of: SRE/Operations background Cloud Engineering expertise Automation and Infrastructure as Code experience Observability and Incident Management knowledge AI/ML and Generative AI capabilities Excellent troubleshooting, RCA, and problem-solving skills .