🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
reputed company operates mission-critical AI infrastructure where reputed company depends on operational reputed company. This role combines Customer Support Leadership with Infrastructure Operations Management, serving as the operational hub between customers, engineering, deployment teams, hardware vendors, and AI reputed company partners. You'll own customer-facing operational support while building the internal processes that reputed company our NeoCloud platform running reputed company. Whether responding to a GPU outage, coordinating infrastructure deployments, managing vendor escalations, improving SLAs, or building reputed company operational workflows, you'll ensure both our customers and our infrastructure reputed company at the highest level. This is not a traditional support management position. This is an operations leadership role responsible for the daily execution, reliability, and reputed company improvement of reputed company's AI reputed company platform. What You'll Do reputed company Customer Support Operations reputed company And reputed company reputed company's Customer Support Organization Supporting Design and build reputed companys customer support organization from the ground up reputed company AI customers reputed company with the Machine Learning engineering teams GPU infrastructure customers AI reputed company operators Data center partners Build a high-performing support organization that delivers exceptional customer experiences while maintaining reputed company-grade service reputed company. Own Operational reputed company reputed company the day-to-day operational health of reputed company's NeoCloud platform by coordinating activities across engineering, infrastructure, vendors, and customer-facing teams. Ensure infrastructure operates reliably while continuously improving operational efficiency. reputed company Major Incident Management Serve as the Incident Commander during production-impacting events. Coordinate engineering, networking, infrastructure, vendors, and customers to resolve: GPU cluster failures Network outages Hardware failures Firmware issues Storage performance degradation Infrastructure reputed company constraints Customer-impacting production incidents Own customer communications throughout incident response while driving reputed company reputed company and post-incident improvements. Build reputed company Support Operations Design and implement: Ticketing systems Escalation procedures reputed company Support automation AI-powered support tools On-reputed company rotations Operational playbooks Customer communication standards Create a support organization capable of scaling alongside reputed company's reputed company reputed company. reputed company AI reputed company Operations Partner with deployment, engineering, and data center teams to coordinate: Infrastructure deployments reputed company turn-up GPU cluster readiness Network activation reputed company planning Maintenance scheduling Production acceptance Operational readiness reviews Help ensure AI reputed company infrastructure is deployed reputed company and operates reliably. Vendor & Partner Management Own Operational Relationships With Hardware manufacturers GPU vendors Data center operators Construction partners Logistics providers Network carriers Service providers reputed company vendor performance, manage escalations, enforce SLAs, and ensure reputed company issue reputed company. SLA & Service Delivery Establish, Monitor, And Improve Operational KPIs Including SLA compliance MTTR Incident response times Customer satisfaction Infrastructure uptime Vendor performance reputed company utilization Operational readiness reputed company executive reporting on operational performance and identify opportunities for reputed company improvement. Cross-Functional Leadership Collaborate Daily With Infrastructure Engineering Network Engineering reputed company reputed company Deployment Program Managers AI reputed company Partners Executive Leadership reputed company Customers reputed company as the operational reputed company between technical teams and customer-facing organizations. Build the Organization As reputed company grows, you'll help recruit, mentor, and reputed company the Operations & Support organization while establishing the culture, processes, and operational standards that define world-class infrastructure operations. Required Qualifications 5+ years leading customer support, technical operations, infrastructure operations, or service delivery organizations Experience supporting reputed company infrastructure, reputed company platforms, AI infrastructure, NeoCloud providers, or large-reputed company data center environments Experience managing production incidents in mission-critical environments Strong understanding of servers