🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
We help the world run better At SAP, we keep it simple: you bring your best to us, and we'll bring out the best in you. We're builders touching over 20 industries and 80% of global commerce, and we need your unique talents to help shape what's next. The work is challenging – but it matters. You'll find a place where you can be yourself, prioritize your wellbeing, and truly belong. What's in it for you? Constant learning, skill growth, great benefits, and a team that wants you to grow and succeed. Meet the team The CF Backing Services SRE team is part of the SAP Business AI Platform organization, responsible for the operational excellence and reliability of business-critical platform services used by thousands of SAP customers globally. Our scope covers services running on both Cloud Foundry and Kubernetes across production landscapes worldwide. We apply SRE principles in practice - not as a philosophy, but as daily engineering work. We own SLOs, we run Chaos Days, we automate toil, and we treat reliability as a feature. We also embrace AI-first tooling as a core part of how we work - from AI-assisted incident response to building and contributing to our own internal SRE tooling. We are looking for an engineer ready to grow into a well-rounded SRE - someone who is curious, technically solid, and motivated to take real ownership of the services they support. What you'll build Live Site Operations You'll participate in hotline and on-call rotation, responding to incidents and SLO violations across supported services You'll investigate production issues with deep technical analysis - log analysis, distributed tracing, Kubernetes debugging, service dependency mapping You'll contribute to RCA creation and post-incident follow-up, including tracking and closing action items You'll participate in Chaos Days and fire drills to proactively test service resilience Reliability Engineering You'll monitor service behavior through SLOs, SLIs, and the 4 Golden Signals - and act on what you find You'll identify and drive improvements to alerting quality - reduce noise, increase signal, eliminate false positives You'll contribute to the team's automation and tooling - scripts, Recommended Actions, runbooks, and internal tools that reduce toil You'll support the onboarding of new services into SRE scope - documentation, monitoring setup, KT sessions, access verification Collaboration and Growth You'll work closely with development teams on reliability topics, improvements, and incident learnings You'll contribute to knowledge sharing within the team - KT sessions, Show&Tells, documentation You'll use and contribute to AI tooling actively - AI-assisted workflows and team-built tools You'll participate in compliance activities and follow internal processes and procedures What We Work With This is not an exhaustive list - it reflects our actual daily environment: Platform and Infrastructure Kubernetes (K8s), Helm, Istio - tools for deployment, traffic management, and debugging ArgoCD - GitOps-based continuous delivery for Kubernetes workloads Cloud Foundry - active stack hosting business-critical platform services AWS, GCP, Azure - multi-cloud landscape coverage Terraform, Concourse, Jenkins - infrastructure as code and CI/CD pipelines Vault, Gardener, Kyma Linux - primary operating environment across all infrastructure Observability Dynatrace - primary observability platform for metrics, traces, and alerting ELK Stack (Elasticsearch, Logstash, Kibana) - log aggregation, search, and analysis Prometheus, Grafana - supplementary monitoring Custom alerting layer for SLO violation tracking Development and Automation Python, Bash - primary scripting languages for automation and tooling GitHub, Jira - version control and task management AI Tooling Joule - SAP internal AI assistant integrated into daily workflows Claude Code, GitHub Copilot - standardized AI-first development workflow Perplexity - used for research and technical investigation Team-built AI skills, workflows, and tools distributed through our internal AI marketplace What you bring Must have: Solid Linux/Unix foundation - comfortable in a terminal, understanding of processes, file systems, and system administration Networking fundamentals - TCP/IP, DNS, HTTP/S, load balancing, basic network troubleshooting Genuine curiosity about how distributed systems fail and how to make them more resilient Ability to communicate clearly and precisely - in incidents, in writing, in cross-team discussions Fluency in English - our team and stakeholders are international A team-first attitude - we share on-call responsibility and we help each other Strong advantage: Experience or solid knowledge of Kubernetes - deploying, debugging, understanding what goes wrong Familiarity with GitOps and tools like ArgoCD Hands-on experience with Python or Bash scripting Experience with log analysis using ELK or similar stacks Familiarity with observability concepts - SLOs, SLIs, alerting, distributed tracing Experience with cloud platforms (AWS, GCP, or Azure) Understanding of CI/CD pipelines and infrastructure as code Experience with databases - particularly PostgreSQL Experience with incident management - on-call rotations, incident response workflows, and RCA processes Strong troubleshooting and problem-solving skills - ability to analyze symptoms, diagnose root causes, and implement effective solutions in distributed systems Mindset we value more than any specific skill: You ask "why is this breaking?" not just "how do I close this ticket?" You look for patterns across incidents, not just fixing individual issues You automate things that repeat - you do not accept permanent toil You use AI tools actively - you see them as a multiplier, not a threat You flag problems early and communicate proactively - no surprises Where you belong Real ownership of business-critical platform services from day one - you will not be a ticket processor A team that invests in knowledge sharing - KT sessions, mentoring, and structured onboarding On-call rotation with special compensation Active AI tooling adoption - you will be working with some of the most advanced AI-assisted SRE workflows in the organization An international, collaborative environment with direct access to the development teams building the services you support Clear career progression path with regular feedback cycles and transparent promotion criteria Education BSc in Computer Science, Engineering, or a related field - or equivalent practical experience. We value demonstrated capability over academic credentials. We know that strong candidates often do not apply because they feel they do not meet every requirement. If you are curious, technically grounded, and genuinely motivated to learn - apply. We have seen people grow into excellent SREs from very different starting points. Bring out your best SAP innovations help more than four hundred thousand customers worldwide work together more efficiently and use business insight more effectively. Originally known for leadership in enterprise resource planning (ERP) software, SAP has evolved to become a market leader in end-to-end business application software and related services for database, analytics, intelligent technologies, and experience management. As a cloud company with two hundred million users and more than one hundred thousand employees worldwide, we are purpose-driven and future-focused, with a highly collaborative team ethic and commitment to personal development. Whether connecting global industries, people, or platforms, we help ensure every
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.