← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

Principal Site Reliability Engineer (New Delhi)

Oracle · Delhi

📅 09/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

See the other 113,540 jobs in India →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Exercises judgment when performing data collection to maintain and optimize operations and reliability. Leverages advanced knowledge to perform incident response and/or maintenance tasks. Provides comprehensive health and performance reporting. Identifies and recommends opportunities for automation. Communicates comprehensive information about services and proactively anticipates and articulates the potential impact of changes. Provides comprehensive support for technology and documents incidents. Conducts advanced experiments with new tools and develops and maintains advanced knowledge of site reliability trends. Career Level - IC4 Key Responsibilities Capacity Ingestion and Management: - Designs and architects infrastructure and/or service according to terms for reliability and functionality. - Forecasts demands for infrastructure and responds to capacity needs, ensuring systems have sufficient resources to handle current and future workloads and identifying resource gaps. - Collaborates with the software development team to develop infrastructures, ensuring features are reliable and scalable according to deployment requirements. - Proactively identifies opportunities for prototyping and drives prototyping initiatives (e.g., testing recent applications or infrastructures, assisting in onboarding) to explore novel approaches. Incident and Service Lifecycle Management: - Exercises judgment when performing data collection, triage, technical analysis, and redirection to maintain and optimize operations and infrastructure reliability. - Takes proactive steps to monitor services, maintain up-to-date knowledge of their performance, and document their condition. - Leverages advanced knowledge to perform incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery). - Provides comprehensive health and performance reporting and takes appropriate actions based on trends in data. - May perform provisioning to support infrastructure, applications, and services. - May experiment with new approaches for and performs decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed. Automation: - Identifies and recommends opportunities for automation and assesses potential benefits to enhance operational efficiency. - Develops and implements design, automation tools, or scripts to provide solutions, gather metrics, monitor, analyze, mitigate, or remediate issues/defects within infrastructures. - Conducts testing on moderately complex automations to ensure they perform tasks correctly and produce expected results. Technical Communication and Guidance: - Writes release notes and/or communicates comprehensive information about the scale, capacity, security, performance attributes, and requirements of services and technology with customers and immediate and related teams. - Proactively anticipates and articulates the potential impact of infrastructure, feature, and tool changes, considering their impact across team operations. - Serves as a resource to team members on what information to communicate and how to communicate. Troubleshooting and Resolution: - Provides comprehensive operational support for technology, serving as a key escalation point for incidents and moderately complex issues arising within Oracle services. - Drives and actively participates in on-call shifts to address issues. - Executes the resolution of technical issues spanning multiple services, applying advanced investigation and debugging techniques to achieve SLOs (service level objectives). - Documents incidents according to reporting methods and performs root cause analyses, capturing essential information for analysis and future reference. - Performs post-mortem procedures to prevent incident reoccurrence. Innovation and Improvement: - Conducts advanced experiments and evaluations of cutting-edge tools and technologies to optimize infrastructure performance and reliability, taking proactive steps to adhere to security standards. - Identifies and seeks opportunities to execute improvements for performance bottlenecks and deployments, ensuring efficient resource usage, speed, and scalability. - Develops and maintains advanced knowledge of site reliability trends, sharing valuable insights and information with senior team members, management, and beyond to promote innovative building, testing, deploying, and running services. - Performs moderately complex analyses and provides clear data on production to drive business development decisions (e.g., design changes). Core Responsibilities Planning & .
Read the rest of the job →
For people searching Engineer

144,883 engineer jobs are open right now

Here's how to pick the right one and stand out in your application.

144.883Jobs
31.687IN
81%EN

That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.

Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.

Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.

When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.

👁 21 have read this
0 comments
Want to comment?

Leave your e-mail to comment, react and follow the posts for your role. It is free.

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →