🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Manager, DevSecOps Production Support Salary: Starting at $170k + bonus Location: Chicago, IL Hybrid: 3 days onsite, 2 days remote *We are unable to provide sponsorship for this role* Qualifications 5+ years of hands-on environment operations, production support, or infrastructure operations experience Experience across: application deployment pipelines, container platform operations, middleware support, incident management, monitoring and observability, configuration management, release engineering, platform operations, or scripting and automation. Formal people management experience Demonstrated experience defining and enforcing SLA frameworks in a tiered support model (L1/L2/L3 or equivalent). Familiarity with financial services or other regulated-industry production environments including knowledge of change governance, audit requirements, and production access controls. Technical skillset Harness (continuous delivery pipelines, deployment verification, rollback automation), Jenkins (CI/CD pipeline management, job configuration, build troubleshooting), GitHub (branching strategies, pull request workflows, pipeline integration). Kubernetes (k8s) pod lifecycle management, namespace operations, log retrieval, resource troubleshooting, and coordination with Platform teams on cluster-level issues. Apache Kafka topic management, consumer group monitoring, lag analysis, and escalation to Platform for broker-level issues. HashiCorp Vault — secrets retrieval, token/lease troubleshooting, policy review, and escalation to Security teams for certificate and secrets rotation. Proficiency in at least two production monitoring toolsets (e.g. Splunk, Dynatrace, Datadog, AppDynamics, Prometheus/Grafana) alert triage, dashboard interpretation, log analysis, and tuning requests. Working knowledge of middleware infrastructure including application servers, messaging brokers, storage integrations, and network-layer dependencies sufficient to triage, gather diagnostics, and route correctly to L3. Responsibilities Lead a team of 6–10 L1 and L2 support engineers Manage team scheduling to ensure full coverage of production support windows including on-call rotations, shift handoffs, and escalation availability for 24×7 support responsibilities. Perform all talent management functions including performance reviews, direct and timely feedback, goal setting, and administrative functions as required. Ensure accurate, complete documentation for every incident — symptoms, steps taken, diagnostics, resolution, and RCA where applicable. Own the runbook library — every novel resolution produces a runbook published to L1 before the incident is closed; coverage gaps are tracked and closed sprint-on-sprint. Lead alert tuning and noise reduction initiatives across monitoring toolsets — on-call engineers are paged for situations requiring human judgement, not system noise. Lead automation and tooling initiatives to reduce toil, accelerate triage, and eliminate manual steps from the support workflow. Define, publish, and enforce SLA targets Monitor SLA compliance in real time; escalate breaches immediately and report trends to leadership on a sprint cadence. Lead L1 and L2 support engineers in all incident response activities including triage, investigation, coordination, resolution, closure, and post-incident reporting. Oversee technical analysis of environment incidents across application deployments, middleware, and platform layers while coordinating response activities with internal engineering, platform, and application development teams. Serve as Tier 3 escalation point for complex incidents beyond L2 capability — triaging, directing, and driving resolution across Platform (k8s, Kafka, TFE), S&I (deployment, middleware, storage, network), Security (Vault, certs, secrets), and App Dev teams.