🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Experience: 13 - 16 Years Location: Bellandur, Bangalore Shift: India / US Shift Role: Principal Architect About the Role We are looking for an experienced AI Observability Principal Architect to define and drive enterprise-wide observability, SRE, AIOps, and intelligent automation strategies. The ideal candidate will have solid expertise in Grafana, OpenTelemetry, SRE practices, AIOps, event management, automation, and infrastructure observability. The role requires an architect who can establish observability standards, implement modern telemetry frameworks, drive AI-driven monitoring initiatives, and enable automated incident management and self-healing across enterprise platforms. Key Responsibilities - Define and implement the enterprise observability strategy, architecture, standards, and best practices. - Design and implement observability solutions covering logs, metrics, and distributed traces using OpenTelemetry. - Develop enterprise dashboards, alerting frameworks, and operational views using Grafana. - Establish SLIs, SLOs, Error Budgets, and Critical User Journeys (CUJs) in collaboration with SRE and application teams. - Lead AIOps and Event Management initiatives including event correlation, anomaly detection, predictive analytics, and alert-noise reduction. - Design and implement incident automation, intelligent routing, remediation, and self-healing capabilities. - Integrate observability platforms with ITSM platforms, preferably ServiceNow ITOM. - Drive automation using PowerShell and Ansible. - Collaborate with SRE, Infrastructure, Application, Cloud, and Service Management teams. - Build executive-level and operational dashboards for service health, performance, availability, and reliability. - Provide technical leadership during major incidents and drive continuous service improvement. - Evaluate and implement AI/ML-driven observability and monitoring solutions. - Support automation initiatives using Power Apps and Power Automate where applicable. Primary Skills - Grafana Observability, Dashboarding & Alerting - OpenTelemetry Logs, Metrics & Distributed Tracing - SRE Practices SLI, SLO, Error Budgets, CUJs - AIOps & Event Management - Splunk / SolarWinds - Azure Observability / Azure Monitoring - Incident Management & Automation - PowerShell - Ansible - AI/ML-driven Monitoring and Predictive Analytics - Strong understanding of Windows/Linux infrastructure, networks, and firewalls Secondary / Nice-to-Have Skills - ServiceNow ITOM - Service Mapping - IntegrationHub / integrations - Identification and Reconciliation Engine (IRE) - AIOps - ServiceNow FSM - ServiceNow ITAM / HAM - Power Apps - Power Automate - Cloud and platform observability - Predictive analytics - AI-driven monitoring and anomaly detection Required Experience - 13+ years of experience across Observability, SRE, IT Operations, Platform Engineering, Monitoring, or AIOps. - Proven experience designing and implementing enterprise observability architectures. - Strong hands-on and architectural expertise in Grafana and OpenTelemetry. - Strong understanding of monitoring, telemetry, alerting, incident management, and automation. - Experience implementing SRE principles, including SLIs, SLOs, error budgets, and reliability engineering. - Experience integrating observability platforms with ITSM tools, preferably ServiceNow. - Strong leadership and stakeholder management skills with the ability to work across infrastructure, application, SRE, and service management teams. Preferred Profile Candidates with experience in the following areas will be preferred: - Enterprise-scale AIOps transformation - AI/ML-based anomaly detection and predictive monitoring - ServiceNow ITOM / AIOps - Automated remediation and self-healing infrastructure - Azure observability - Large-scale Grafana and OpenTelemetry implementations - Major Incident Management and Continuous Service Improvement .
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.