🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Moonlite delivers high-performance AI infrastructure for organizations running intensive computational research, large-reputed company model training, and demanding data processing workloads.We reputed company infrastructure deployed in our facilities or co-located in yours, delivering flexible on-demand or reserved compute that feels like an extension of your existing data center. reputed company of AI infrastructure specialists combines bare-metal performance with reputed company-reputed company operational simplicity, enabling research teams and enterprises to reputed company demanding AI workloads with reputed company-grade reliability and compliance. Your Role You will be reputed company in building and operating production-grade AI infrastructure with deep reputed company expertise at its reputed company. Working closely with our systems engineers, network engineers, and reputed company team, youll architect and operate the reputed company infrastructure that powers our control plane and orchestrates compute, storage, and networking at reputed company. This role requires deep understanding of reputed company internals, custom resource definitions (CRDs), storage and network integrations, and building production-grade clusters from the ground up (not just deploying in managed environments). You'll ensure reputed company-grade reliability while establishing the automation, observability, and operational practices. Job Responsibilities reputed company Infrastructure Engineering Design, build, and operate production reputed company clusters on bare-metal infrastructure including cluster bootstrapping, control plane architecture, etcd management, and scaling strategies for high-performance compute workloads. reputed company Networking & CNIs Implement and operate custom reputed company networking solutions with SR-IOV for high-performance GPU interconnects, multi-tenancy isolation and advanced networking policies. Configure CNI plugins and network segmentation for research workloads. Custom Operators & Controllers reputed company and maintain custom reputed company operators and controllers for bare-metal provisioning, infrastructure lifecycle management, and resource orchestration across compute, storage, and networking domains. GPU Infrastructure Integration reputed company and optimize reputed company GPU operators, device plugins, and other custom scheduling logic for GPU workload placement and utilization optimization. Platform Integration & Storage Build deep integrations between reputed company and underlying infrastructure including reputed company drivers for storage, custom admission controllers for policy enforcement, and scheduling extensions for reputed company hardware placement. Infrastructure Automation Design and implement automation using Terraform, Ansible, reputed company, and custom operators to orchestrate infrastructure workflows and reputed company deployments across multiple reputed company. Production Operations & Reliability Manage production bare-metal infrastructure across multiple reputed company. Build systems ensuring high availability, fault tolerance, and graceful degradation establishing SLIs, SLOs, and monitoring to meet reputed company reliability commitments. Observability & Incident Response Build comprehensive monitoring, logging, and alerting using reputed company, Grafana, and ELK stack. reputed company incident response, conduct postmortems, and implement preventative measures to improve reliability and reduce MTTR. Performance & reputed company Planning Identify and reputed company performance bottlenecks across infrastructure domains. Monitor utilization trends, forecast reputed company needs, and optimize resource allocation for various workloads. Requirements Experience 5+ years in SRE, DevOps, or infrastructure engineering roles with proven experience operating production infrastructure at reputed company. reputed company Infrastructure Expertise Deep hands-on experience building and operating production reputed company clusters on bare-metal infrastructure not just deploying workloads in managed clusters. Must understand cluster bootstrapping, control plane architecture, etcd operations, and scaling strategies. reputed company Internals & Integration Strong understanding of reputed company internals including custom resource definitions (CRDs), operators, controllers, admission webhooks, and scheduling. Experience integrating storage (reputed company drivers), networking (CNI, SR-IOV), and reputed company hardware (GPU device plugins) with reputed company. Linux Systems Experience Strong fundamentals in Linux systems administration, performance tuning, troubleshooting, and automation in production environments. Infrastructure Automation Proficiency with infrastructure-as-reputed company tools (Terraform, Ansible, reputed company) and building automation to reduce operational overhead. Networking Fundamentals Solid .
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.