← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

Principal Site Reliability Engineer

Oracle · Nashville, TN

📅 20/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

View and apply on WorkMundi →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Job Description The successful candidate will serve as a senior technical authority, establish reliability standards, guide complex technical decisions, and lead improvements that reduce operational risk and manual effort. This individual must be comfortable moving between architecture and hands-on execution, including accessing deployed hosts, troubleshooting failed services, reviewing logs, correcting configurations, and validating production changes. Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Exercises judgment when performing data collection to maintain and optimize operations and reliability. Leverages advanced knowledge to perform incident response and/or maintenance tasks. Provides comprehensive health and performance reporting. Identifies and recommends opportunities for automation. Communicates comprehensive information about services and proactively anticipates and articulates the potential impact of changes. Provides comprehensive support for technology and documents incidents. Conducts advanced experiments with new tools and develops and maintains advanced knowledge of site reliability trends. Responsibilities Key Responsibilities Design and architect reliable, secure, scalable, and maintainable infrastructure and services. Take proactive steps to ensure solutions meet defined reliability and functionality requirements. Establish technical direction, engineering standards, and operational best practices across complex infrastructure and application environments. Identify system dependencies, operational risks, capacity constraints, performance issues, and potential failure points before they affect service. Translate business, client, security, and application requirements into practical infrastructure and reliability solutions. Lead the installation, configuration, deployment, and validation of applications across Windows Server and Linux environments. Oversee structured builds and deployments using runbooks, scripts, readiness assessments, change controls, and post-deployment validation. Troubleshoot complex operating system, service, application, installation, patching, permissions, certificate, and connectivity issues. Define and improve monitoring, alerting, logging, observability, capacity planning, and service-health practices. Develop and promote automation that reduces manual effort, improves consistency, and lowers operational risk. Own the administration, governance, and continuous improvement of an enterprise infrastructure automation platform, including maintaining automation standards, enabling engineering teams, and driving adoption of scalable, repeatable infrastructure management practices. Lead operating system, middleware, and application patching initiatives, including change planning, rollback preparation, execution, and validation. Direct major incident response, root cause analysis, corrective-action planning, and prevention of recurring failures. Partner with cybersecurity teams on vulnerability remediation, system hardening, STIG compliance, and other security-driven changes. Evaluate emerging technologies and recommend solutions that improve reliability, resilience, security, and operational efficiency. Create and maintain technical standards, architecture documentation, runbooks, deployment procedures, and troubleshooting guides. Provide technical leadership, mentorship, and design guidance to engineers across multiple teams. Communicate technical risks, dependencies, decisions, and recommendations clearly to leadership and stakeholders. Core Skills And Qualifications Technical Leadership and Architecture Extensive experience in site reliability engineering, systems engineering, infrastructure architecture, production operations, or application hosting. Demonstrated ability to design and support highly available, resilient, and secure enterprise systems. Experience leading complex technical initiatives across engineering, operations, security, networking, and application teams. Ability to make sound architectural decisions, evaluate tradeoffs, and communicate recommendations to technical and nontechnical stakeholders. Experience defining engineering standards, operational controls, and reliability practices. Windows and Linux System Administration Advanced, hands-on experience administering Windows Server and/or Linux systems. Ability to access deployed hosts and perform post-deployment configuration, troubleshooting, and validation. Experience installing, configuring, and validating applications in Windows Server and Linux environments. Ability to resolve operating-system-level, service-level, and application-level issues. Strong knowledge of system services, permissions, configuration files, logs, processes, and resource utilization. Manual Build And Deployment Experience Experience leading structured build and deployment activities using runbooks, deployment guides, scripts, and technical procedures. Ability to execute and troubleshoot scripts, validate outputs, and resolve build or configuration issues. Experience with build handoffs, environment-readiness assessments, deployment validation, and post-build verification. Ability to identify process gaps, document exceptions, and improve deployment procedures. Experience managing complex or high-risk production changes. Troubleshooting and Operational Support Ability to investigate complex service failures, installation errors, patching failures, application startup problems, permissions issues, and connectivity incidents. Experience reviewing logs, event viewers, service status, configuration files, ports, certificates, and access controls. Strong analytical and problem-solving skills, with the ability to isolate root causes and implement sustainable solutions. Experience leading major incident response and coordinating technical teams during business-critical outages. Ability to document symptoms, findings, impact, corrective actions, and recommended next steps clearly. Extensive experience supporting production or other mission-critical environments. Scripting and Automation Advanced hands-on experience with one or more of the following: PowerShell Bash Python Ansible Chef Candidates should be able to create, modify, validate, and troubleshoot scripts and automation workflows. Experience identifying automation opportunities and establishing safe, repeatable operational processes is essential. Experience administering or owning an enterprise automation platform for infrastructure management is strongly preferred. Candidates should be comfortable serving as the technical owner of automation capabilities, establishing operational standards, supporting users, and driving ongoing platform adoption and improvement. Patching and Software Maintenance Experience planning and executing operating system, middleware, and application patching. Ability to troubleshoot patch failures, compatibility issues, and post-patch application problems. Strong understanding of maintenance windows, change control, rollback planning, risk assessment, and post-change validation. Experience coordinating patching and remediation activities across application, infrastructure, cybersecurity, and client teams. Cloud and Oracle Cloud Infrastructure Strong understanding of cloud-hosted and hybrid infrastructure. Experience with Oracle Cloud Infrastructure or another major cloud platform. Knowledge of cloud compute, storage, networking, identity, access management, load balancing, and environment provisioning. Experience designing or supporting reliable, secure, and scalable cloud environments. Familiarity with infrastructure-as-code and configuration-management pra
Read the rest of the job →
For people searching Engineer

144,883 engineer jobs are open right now

Here's how to pick the right one and stand out in your application.

144.883Jobs
31.687IN
81%EN

That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.

Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.

Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.

When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.

👁 21 have read this
0 comments
Want to comment?

Leave your e-mail to comment, react and follow the posts for your role. It is free.

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →