🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
We are seeking an AI Infrastructure Engineer to help build, deploy, and operate the LLM serving infrastructure underpinning our AI/ML platform. This role focuses on implementation, automation, and optimization of production inference systems, tuning model serving runtimes for performance and scale across NVIDIA and AMD GPU fleets, working under the technical direction of senior engineering leadership and the platform's established architectural patterns. You'll ship well-engineered, well-tested infrastructure changes and grow your depth in GPU-backed workloads, distributed model serving, observability, and continuous delivery. You'll work directly with senior engineers on real production systems, receive code and design review on everything you ship, and have a clear path to expanded scope and ownership as your experience deepens. What You'll Do Deploy, tune, and optimize high-performance LLM inference pipelines on GPU infrastructure, improving throughput, latency, and cost efficiency within established design patterns. Analyze, profile, and optimize model serving workloads across inference frameworks such as vLLM, SGLang, and TensorRT-LLM, and across different model families and hardware architectures. Build and operate scalable, production-grade API services for model inference, including request routing, multi-tenant isolation, usage metering, and observability. Develop benchmarking harnesses, monitoring infrastructure, and automation tooling that make serving performance measurable and reproducible. Scale inference workloads across multi-GPU, multi-node environments spanning NVIDIA and AMD accelerators. Evaluate, prototype, and integrate model fine-tuning workflows and frameworks. Collaborate closely with engineering and product teams to align infrastructure capabilities with customer-facing services. Investigate and resolve issues across the stack, including container, node, network, and accelerator-level problems, escalating appropriately when scope exceeds the role. Write clear documentation, including runbooks, internal references, and design notes for the changes you ship. Participate in code and design reviews, both as author and reviewer, and incorporate feedback from senior engineers into your work. Required Qualifications Bachelor's degree in Computer Science, Computer Engineering, Applied Math, or Data Science, plus three (3) years relevant work experience or equivalent combination of education and relevant experience. Professional software engineering experience, with at least some of it touching ML systems, GPU workloads, or high-performance backend services. Working knowledge of Kubernetes in a production context, including writing and debugging manifests, understanding core resource types, and operating production workloads. Hands-on experience serving or deploying LLMs — you've run vLLM, SGLang, TGI, TensorRT-LLM, or similar. Comfort working in a Linux environment and with standard developer tooling, including Git-based workflows. Familiarity with CI/CD systems and the basic mechanics of automated build, test, and deployment pipelines. Strong proficiency in Python with familiarity at least one programming or scripting language used for infrastructure work (Go, Rust, C++ or Bash). Preferred Qualifications Experience building or operating retrieval-augmented generation (RAG) pipelines, including vector databases, embedding models, and retrieval serving at scale. AMD/ROCm experience. Experience cleaning and curating datasets for LLM training and fine tuning. Fine-tuning experience of any depth: LoRA/QLoRA, full fine-tunes, dataset curation, or evaluation design. Experience with usage metering, billing systems, or multi-tenant API platforms. Experience instrumenting services and consuming observability data, including writing Prometheus queries, building Grafana dashboards, or working with distributed traces. Experience with HPC batch schedulers and MPI based workloads. Experience with alternative compute architectures for inference (RISC-V, FPGA, ASIC). Why Join Hoonify You'll have a direct line to leadership and genuine influence over the company's growth trajectory. This is a rare opportunity to build a cutting-edge multi-cloud computational platform at a company doing meaningful work in AI — with the autonomy and resources to make it your own. About Our Team Hoonify delivers secure, sovereign AI infrastructure designed for the next generation of inference workloads. Powered by TurbOS , our platform enables organizations and NeoCloud/data center operators to transform CPU/GPU infrastructure into production-ready AI environments—supporting local LLMs, agentic copilots, RAG, and embeddings. We empower teams with robust model lifecycle management, multi-tenant controls, usage metering, and fully auditable operations. Hoonify is an equal opportunity employer. We welcome applicants from all backgrounds and are committed to building a diverse and inclusive team. Must be eligible to obtain and maintain a US government security clearance.
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.