← WorkMundi · 1M+ jobs from around the world, liveSign inCreate free account

Machine Learning Performance Engineer

Long Ridge Partners · New York, NY

📅 20/08/2026
🔔 Alert me about jobs like this
No password, no sign-up. Just the email — and you can leave the list anytime.
🔓 Apply — free →
Opens this job on WorkMundi. The account is free and takes under a minute.

See the other 152,342 jobs in United States →

🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Machine Learning Performance Engineer (Inference) High-Frequency Trading Firm Compensation: $600,000-1.5 million total About the Opportunity A leading high frequency trading firm is hiring a Machine Learning Performance Engineer to sit at the intersection of quantitative research and high-performance production systems. In his role, you'll architect inference pipelines that operate at the physical limits of hardware, driving the speed, efficiency, and reliability of ML inference so predictive models consistently achieve microsecond-level latency. GPU usage across the firm's trading teams has grown roughly 100x in the past year as deep learning has moved from a supporting signal to the core of how strategies are built. That growth has outpaced the decision-making around it. Strategies get pushed onto GPUs by default, without anyone systematically asking whether GPU is the right target at all. This role owns that question end to end: benchmark the workload across CPU, GPU, and FPGA, decide the architecture on evidence, then optimize and deploy against it. You will also have the chance to revisit existing models that never reached production, some of which stalled for hardware or deployment reasons, and run them through different environments to determine where they belong. What You'll Do Benchmarking & Strategy Lead the technical evaluation of inference platforms across CPUs, GPUs, and FPGAs to guide infrastructure deployment decisions Benchmark trading workloads across architectures before compute is committed, and identify where performance gains actually come from — code-level or hardware-level System Architecture Optimization Analyze and enhance execution across deep memory hierarchies to maximize resource utilization and parallel processing Assess and resolve memory subsystem and interconnect bottlenecks across the end-to-end inference lifecycle Infrastructure & Deployment Feasibility Work with Infrastructure teams to understand the thermal, power, and operational constraints of hardware platforms, and design inference strategies for latency-critical trading strategies that fit within those envelopes Consider the interaction between trading workloads, compute requirements, hardware selection, and fleet utilization GPU Kernel Development Develop highly optimized kernels and integrate specialized performance libraries to extract maximum computational throughput from the underlying silicon Model Optimization & Deployment Implement advanced model reduction techniques — quantization, pruning, distillation — to ensure compact memory footprints and numerical stability Prioritize optimization for low-latency, event-level inference workloads that meet real-time trading requirements Cross-Functional Collaboration Partner closely with ML Researchers, HPC Engineers, FPGA Engineers, and Datacenter Engineers to bring target deployments to production What We're Looking For 2+ years optimizing deep learning inference in latency-sensitive or high-throughput production environments, in any domain ML frameworks: deep expertise in lower-level ML framework development (PyTorch/JAX), paired with strong Python/C++ skills and a thorough understanding of mixed-precision computation Kernel development and tooling: proven experience building custom GPU kernels, with deep familiarity with optimization libraries and compilers (Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS) and profiling tools (Nsight Systems, Nsight Compute) GPU architecture: deep expertise in GPU microarchitecture, including SM execution, warp scheduling, and full memory hierarchy optimization from registers to HBM Cross-architecture benchmarking: a rigorous, data-driven track record evaluating inference performance across heterogeneous compute architectures Prior experience in financial trading is not required. Nice to Have Practical experience targeting and optimizing inference workloads on specialized hardware ecosystems, including FPGAs and ASICs Why Join? This is a role with genuine decision-making scope. Rather than optimizing code for whatever hardware happens to be available, you will determine which hardware the workload should run on in the first place, prove it with data, and then build for it. That combination of architectural judgment and hands-on kernel, and the results are measurable in production almost immediately. You'll work on inference at microsecond latency, where the constraints are physical rather than theoretical, and where memory hierarchy, interconnect behavior, thermal envelopes, and fleet utilization all shape the answer. Benefits include generous paid time off, regional savings and financial wellness plans, hybrid working options, free breakfast, lunch, and snacks daily, in-office wellness experiences and reimbursement for select wellness expenses, company-sponsored sports teams and fitness events, volunteer and charitable giving opportunities, regular social events, and ongoing workshops and learning opportunities. The culture is collaborative and low on hierarchy, smart, driven people, an open-plan workspace, casual dress, and an environment where the best idea wins. Equal opportunity employer.
Read the rest of the job →
For people searching Engineer

144,883 engineer jobs are open right now

Here's how to pick the right one and stand out in your application.

144.883Jobs
31.687IN
81%EN

That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.

Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.

Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.

When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.

👁 21 have read this
0 comments
Want to comment?

Leave your e-mail to comment, react and follow the posts for your role. It is free.

Similar jobs

Job on WorkMundi — the world's largest job board. See more jobs from every continent, updated live.

📢
🎁

Before you apply, rehearse this interview.

Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card.

I want my training →