🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Job Title: Site Reliability Engineer (SRE) L2 Support (Apigee + Microservices) Job Summary We are seeking an analytical and technically skilled SRE L2 Support Engineer with 2 to 3 years of experience specializing in API management and microservices infrastructure. In this role, you will act as the escalation point for the L1 monitoring team, taking ownership of deep technical triage, debugging complex application errors, and optimizing API gateways. You will focus on maintaining high uptime for critical banking application platforms, identifying systemic errors, and implementing permanent fixes or automated workarounds within a rapid-paced cloud workplace. Key Responsibilities L2 Triage & Advanced Incident Management Escalation Ownership: Act as the direct L2 technical escalation point for complex infrastructure, application, and API connectivity issues routed from L1. Deep-Dive Debugging: Analyze verbose log traces, evaluate application stack metrics, and dissect error signatures to quickly pinpoint core system failures. Crisis Collaboration: Partner closely with L3 engineering, backend application development, and DevOps teams during major (P1/P2) production incidents to accelerate system recovery. Apigee & Microservices Management Apigee API Administration: Troubleshoot API proxy execution issues, configure security policies (OAuth, API keys, spike arrests), resolve SSL handshake errors, and debug payload routing failures. Microservices Support: Monitor, debug, and trace distributed applications across complex microservice dependencies using distributed tracing logs. Kubernetes Orchestration: Manage cluster health, troubleshoot pod eviction or crash loops, view container resource limits, and run diagnostic deployments utilizing advanced kubectl commands. Observability, Reliability & Automated Recovery Telemetry Dashboards: Navigate and customize production monitoring workflows within Datadog, Dynatrace, Prometheus, and Grafana to narrow down systemic bottle .