🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
About the role Leads complex, data center hardware support services to ensure resilience, reliability, and availability at scale. Designs, governs, and signs off on advanced change activities and site augmentations, driving standardization of SOPs/MOPs across regions. Serves as the highest operational escalation point for incidents, leading cross-functional triage, root-cause analysis, and long-term corrective actions. Partners with engineering, capacity, networking, and security teams to anticipate risks, implement architectural improvements, and automate repeatable workflows. Mentors and develops team mates, uplifts best practices, and influences policy and process changes that improve performance across the group and organization. Key responsibilities Oversees lifecycle management for servers and components; sets standards for diagnosis, replacement, and performance validation without impacting critical workloads Coaches teams on advanced power distribution practices and complex hardware interventions; validates that work meets reliability, safety, and SLA objectives Leads the planning, risk assessment, and execution governance for large-scale network builds and migrations; validates architecture, cabling, and end-to-end connectivity Defines acceptance criteria and sign-off gates for network changes; ensures interoperability and error-free integration with upstream and downstream systems Owns the escalation queue and trend analysis; directs complex investigations, establishes fix-forward plans, and ensures high-quality documentation of resolutions within SLAs Shapes the ticket taxonomy, triage playbooks, and automation triggers to improve response consistency and speed Governs adherence to SOPs/MOPs and site rules for all high-risk work; ensures alignment with physical security procedures, local regulations, and audit requirements; verifies functional testing of security controls Ensures comprehensive documentation, audit trails, and change records; prepares for and leads compliance reviews and remediations Leads post-incident reviews and enterprise-level RCAs; socializes learnings and implements durable design/process changes that reduce risk across the fleet Mentors technicians and specialists; sets expectations for execution quality, safety, and customer focus; builds capability through training and coaching About you 6 years of experience in client-facing, direct customer support and service, data center, or related field experience, OR High School Diploma, General Educational Development (GED), or equivalent AND 5 years of experience in client-facing, direct customer support and service, data center, or related field experience, OR Associate Degree in Computer Science, Information Technology, Data Center Operations or related field AND 4 years of experience in client-facing, direct customer support and service, data center, or related field experience 4 years of experience in IT infrastructure, data center services, or server administration