← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

SRE | Kubernetes | Datadog | AWS | Python | Remote India | Remote/WFH

Remote_WFH_INDIA · All India

🌐 Remote📅 24/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

View and apply on WorkMundi →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Senior Site Reliability Engineer | Datadog | Kubernetes | AWS | Python | CI/CD | Remote India | JST Overlap We're AgileEngine! AgileEngine is an award-winning software engineering and AI services company partnering with Fortune 500 companies to build world-class digital products. Youll work extensively with Datadog, Kubernetes, AWS, Python, APIs, and CI/CD pipelines, with a strong focus on building reliable and scalable observability solutions. Location: Remote / WFH India Timezone: JST timezone overlap is mandatory Experience: 4+ years Client: Indeed --- Why You Should Apply Work on a high-impact platform reliability initiative for Indeed Build and maintain enterprise-grade Datadog observability solutions Work with Kubernetes-based microservices environments Design dashboards, alerts, APM, metrics, logging, and tracing Integrate observability into AWS and CI/CD pipelines Automate monitoring and operational tasks using Python Work across software engineering and site reliability engineering Drive platform reliability, scalability, performance, and continuous improvement Take ownership of platform modernization and maintenance initiatives --- Must-Have Skills 4+ years of professional software engineering / SRE experience Strong proficiency in Python, JavaScript (Node.js), or Java Strong hands-on experience with Kubernetes deployment, operations, and monitoring Hands-on experience with Datadog or similar observability platforms such as Prometheus/Grafana Experience configuring Datadog dashboards, alerts, APM, metrics, logging, and tracing Experience monitoring containerized and microservices-based applications Hands-on experience with AWS Experience integrating observability tools into cloud environments Experience integrating observability into CI/CD pipelines Strong experience with API integrations designing, consuming, and integrating APIs Experience automating monitoring and operational tasks using scripting, preferably Python Experience installing and configuring Datadog agents and integrations Understanding of secure configuration, API keys, user roles, and access controls Upper-intermediate English communication skills JST timezone overlap Mandatory --- What You'll Do Support platform reliability, monitoring, and continuous improvement across internal systems Work extensively in Kubernetes-based environments Build and maintain Datadog dashboards, alerts, APM, metrics, logs, and traces Monitor containerized and microservices-based applications Integrate observability solutions with AWS environments Integrate monitoring and observability into CI/CD pipelines Install and configure Datadog agents and integrations Manage API keys, secure configurations, user roles, and access controls Automate monitoring and operational activities using Python Lead maintenance initiatives and platform improvements Improve system reliability, scalability, and performance Proactively identify opportunities to modernize and improve platform operations --- Nice to Have Experience owning and operating an internal engineering platform Demonstrated ownership of reliability, scalability, and performance Experience proactively leading maintenance and platform improvement initiatives Familiarity with Golang Experience with New Relic, Dynatrace, Elastic, or Splunk Observability Strong experience with distributed and microservices architectures --- Priority Will Be Given To Candidates With Strong Python experience Strong Kubernetes experience Hands-on Datadog experience Strong AWS experience CI/CD integration experience Observability experience across APM, metrics, logging, tracing, dashboards, and alerts Experience monitoring containerized / microservices applications Strong API integration experience Experience automating operational tasks using Python Proven ownership of reliability, scalability, and performance JST timezone overlap --- Tech Stack Datadog Kubernetes AWS Python CI/CD Observability APM Monitoring Metrics Logging Tracing Microservices APIs Docker/Containers Cloud Site Reliability Engineering --- Important Timezone Requirement Candidates must be able to provide the required JST timezone overlap for this role. Please apply only if you are comfortable working with the required JST overlap. --- Hiring Process Step 1: Technical Assessment Step 2: Video Interview Step 3: Recruiter Discussion Step 4: Technical Interview Step 5: Offer Please complete all LaunchPod steps promptly to avoid delays in the hiring process. --- To Apply, DM me with: 1. Email ID 2. Total Experience 3. SRE Experience 4. Python Experience 5. Kubernetes Experience 6. Datadog Experience 7. AWS Experience 8. CI/CD Experience 9. Observability Experience 10. API Integration Experience 11. Microservices / Container Experience 12. Othe
Read the rest of the job →

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →