🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
What Youll Be Doing:Partner with NVIDIAs software, research, architecture, and product teams to align technical requirements and strategic priorities, fostering the AI ecosystem on RTX and DGX PCs.Build and optimize the local AI inference stack for RTX, RTX Pro, and DGX GPUs, with a focus on performance, stability, and scalability across diverse hardware architectures.Design and develop modern inference runtimes and execution stacks using frameworks such as llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX, supporting LLM, vision-language, TTS, ASR, and diffusion-based AI workloads.Perform end-to-end optimization of AI models, data pipelines, and inference runtimes to maximize performance on current and next-generation GPU architectures. Apply model optimization techniques, including quantization, pruning, sparsity, and distillation, to enable efficient deployment of large models on local and edge devices.Conduct system-level debugging, performance tuning, and performance-accuracy trade-off analysis; develop infrastructure for performance and accuracy sweeps; analyse results to identify gaps and drive fixes; and establish engineering guidelines to accelerate bring-up and ensure production readiness of new models and inference backends.What we need to see:5+ Years of experience with Bachelors, Masters, or PhD in Computer Science, Software Engineering, Mathematics, or a related field, or equivalent experience.Excellent C++ programming and debugging skills, with a strong foundation in data structures, algorithms, and machine learning.Proven experience developing and optimizing AI inference pipelines and applications using ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, Windows ML, DXCGC, and TensorRT.Deep understanding of inference backends and runtime internals, including scheduling, memory management, KV-cache behaviour, graph execution, quantization, and hardware-aware optimization.Strong analytical and problem-solving skills, with the ability to manage multiple priorities effectively in a fast-paced environment.Excellent written and verbal communication skills, enabling effective collaboration across engineering teams and management.Ways to stand out from the crowd:Understanding of modern machine learning, deep neural network, and generative AI techniques, with relevant contributions to major open-source projects.Consistent track record of delivering end-to-end products in multinational companies with geographically distributed teams.Proficiency in low-level system and GPU programming, CUDA, and the development of high-performance systems.Contributions to open-source inference runtimes, model tooling, or performance infrastructure.Hands-on experience building applications using frameworks and APIs such as llama.cpp, PyTorch, TensorRT, Vulkan, DirectX, and vLLM.We're a top employer recognized for innovation, growth, and a commitment to diversity as an equal-opportunity workplace. We offer competitive salaries, a generous benefits package, and the opportunity to work alongside some of the technology industry's most talented and forward-thinking professionals. As our engineering teams continue to grow rapidly, we're looking for creative, self-driven engineers with a passion for technology to join us. .
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.