🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
BCforward is seeking a highly motivated and experienced Observability Operations Engineer Note: Candidate must be local to Phoenix, Arizona. Job Title: Observability Operations Engineer Job Location: Phoenix, AZ Hybrid Duration: Long-term Pay Rate: $60/hr W2 and $68 CTC Must Haves: Dynatrace and Splunk, as well as OpenSearch/elastisearch and OpenTelemetry. Kubernetes AI/ ML, Grafana and Sahara Automation using AI Role is 20% automating at 80% operations. 3 days a week. Standard shifts are either 9:00 AM to 6:00 PM or 10:00 AM to 6:30 PM to provide coverage. Candidates also must be willing to work 1 weekend day every 2-3 weeks. Job Description: We are seeking a highly skilled Senior Observability Operations Engineer to manage and enhance our enterprise observability platform. The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and cloud-native observability solutions. Experience leveraging AI/ML and Generative AI to improve observability, automate operations, and accelerate incident resolution is highly desirable. The role is responsible for ensuring high availability, scalability, operational excellence, and continuous improvement of enterprise monitoring and logging platforms supporting mission-critical applications. Key Responsibilities Administer and optimize enterprise observability platforms including Dynatrace, Splunk, and OpenSearch/Elasticsearch. Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions. Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning. Configure Dynatrace OneAgent, ActiveGate, Synthetic Monitoring, Real User Monitoring (RUM), Digital Experience Monitoring (DEM), Davis AI, and Application Performance Monitoring (APM). Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI. Develop dashboards, alerts, reports, and executive operational metrics. Support Linux-based infrastructure and Kubernetes environments (Docker/OpenShift/Rancher preferred). Implement observability best practices using OpenTelemetry, distributed tracing, metrics, logs, and events. Perform root cause analysis for production incidents using observability platforms. Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams. Automate operational tasks using Python, Shell scripting, REST APIs, Terraform, or Ansible. Participate in incident, problem, change, and release management processes. Drive platform upgrades, patching, security compliance, and operational governance. Improve platform reliability through automation, self-healing, and AI-assisted operations. Required Technical Skills Observability Platforms Dynatrace Administration Splunk Enterprise Administration OpenSearch Administration Elasticsearch Administration Grafana Prometheus Kibana Jaeger OpenTelemetry Kafka (preferred) Infrastructure Linux Administration Kubernetes Docker OpenShift or Rancher Networking (TCP/IP, DNS, Load Balancers, Firewalls) System Administration Cloud & DevOps AWS, Azure, or GCP CI/CD pipelines Git Terraform Ansible REST APIs Scripting Python Bash/Shell PowerShell (preferred) AI & Automation Skills (Preferred) Experience using Generative AI (ChatGPT, GitHub Copilot, Amazon Q, Microsoft Copilot, or similar) to improve operational efficiency. Knowledge of AIOps platforms and AI-driven observability. Experience with Dynatrace Davis AI for anomaly detection and root cause analysis. Understanding of machine learning concepts for predictive monitoring and intelligent alerting. Experience building AI-assisted operational runbooks and troubleshooting workflows. Knowledge of Retrieval-Augmented Generation (RAG), vector databases, embeddings, and AI-powered knowledge search is a plus. Experience integrating AI with observability platforms using APIs. Familiarity with LLMs, prompt engineering, and AI-assisted automation. Experience using Python with AI frameworks (LangChain, LangGraph, OpenAI APIs, or similar) is desirable. Exposure to AI-driven incident summarization, log analysis, and automated ticket enrichment. Required Qualifications Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience. 6–10+ years of IT infrastructure or observability operations experience. 4+ years administering Dynatrace, Splunk, OpenSearch, or Elasticsearch. Strong Linux system administration experience. Experience supporting enterprise-scale production environments. Strong troubleshooting and analytical skills. Excellent communication and stakeholder management skills. Preferred Certifications Dynatrace Associate or Professional Certification Splunk Enterprise Certified Administrator Elastic Certified Engineer Kubernetes (CKA/CKAD) AWS/Azure/GCP Certification ITIL Foundation AI/ML or Generative AI certification (preferred) Soft Skills Strong ownership and accountability Excellent problem-solving and analytical thinking Ability to work independently with minimal supervision Strong collaboration across cross-functional teams Continuous learning mindset Ability to thrive in fast-paced production environments
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.