🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Company Description TensorMem Inc is a fast-paced deep-tech startup at the intersection of Cloud Scale Storage and Artificial Intelligence. Our mission is to build the next generation of distributed computing and infrastructure solutions for AI workflows. Our software-defined, inference-native context-memory orchestration platform for AI, focuses on solving performance and cost challenges in large-scale inference. As AI inference grows with long-context models, RAG pipelines, agentic systems, and copilots, TensorMem addresses memory walls, KV cache pressure, and inefficient data movement that limit throughput and efficiency. The platform treats inference state, context, and memory as first-class resources across the AI memory hierarchy, enabling predictable performance, fewer GPU stalls, and better infrastructure ROI. We are looking for an extraordinary Principal Engineer to join our core engineering team to architect, design, and build TensorMem's next-generation distributed inference systems, leveraging modern AI-assisted development practices. Role and Responsibilities As a Principal Engineer for Distributed Inference Systems, you will serve as a primary technical authority driving the architecture, design, and performance strategy for TensorMem's high-scale distributed storage and compute platforms. You will solve complex systems-level engineering challenges to deliver ultra-low latency and high-throughput AI inference across heterogeneous hardware clusters. Your primary responsibilities will include: Architecting, designing, and building production-grade, large-scale distributed storage and compute systems tailored for distributed AI inference workloads.Designing and implementing distributed algorithms, fault-tolerant consensus protocols, data partitioning, and replication strategies across multi-node clusters.Applying deep Linux systems expertise (NUMA, PCIe, NVMe, RDMA, custom memory management) to optimize system-level data movement and execution paths.Integrating distributed filesystems, object stores, or databases seamlessly with GPU/XPU acceleration layers to maximize inference throughput and minimize latency.Collaborating with AI infrastructure teams to optimize modern inference engines (e.g., vLLM, TensorRT-LLM, Triton) for GPU-accelerated environments.Setting engineering standards, conducting high-impact architecture reviews, and mentoring senior engineers across the organization. Qualifications Education & Experience Bachelors, Master's, or Ph.D. degree in Computer Science, Computer Engineering, Electrical Engineering, or a closely related technical field from a reputed institution.1520 years of hands-on, professional experience designing, building, and scaling large-scale distributed compute, storage, or cloud infrastructure systems.Proven track record of architectural leadership and technical strategy in high-performance or deep-tech startup environments. Technical Skills (Must-have) Exceptional mastery of Computer Science fundamentals, including advanced data structures, core system algorithms, and distributed systems algorithms/protocols (e.g., Raft, Paxos, gossip protocols, distributed locking).Deep, expert-level understanding of Linux systems architecture, including NUMA architectures, PCIe bus topologies, NVMe storage, RDMA network interfaces, and low-level memory management.Excellent programming skills inC++,Go, andPython.Significant hands-on experience designing and building distributed filesystems, distributed databases, or object storage systems (e.g., POSIX, Ceph, Lustre, NVMe-oF, S3-compatible systems).Good understanding of modern AI inference engines and how execution, memory layout and sharing, and tensor operations work on GPU hardware accelerators. Soft Skills Proven technical leadership with the ability to articulate architectural vision, mentor engineering talent, and influence strategic product roadmaps.Excellent problem-solving skills, rigorous engineering discipline, and strong attention to detail.Ability to work effectively in a fast-paced, high-impact, and collaborative startup environment.A strong eagerness to integrate AI tools (e.g., GitHub Copilot, integrated IDE features) into the daily engineering workflow to maximize efficiency. .
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.