Senior ML Engineer, Serving & Optimization
Join one of the most sophisticated trading firms in the world as they build a new machine learning initiative from the ground up. The team is currently just three people, including engineers from two competitors and a researcher from a top FAANG AI lab, and they're looking to add a few key hires who will help define the future of ML infrastructure at the firm.
This particular hire will focus on model serving, inference optimization, and deploying large-scale models onto GPU infrastructure. The team is already operating at scale with over 1,000 GPUs today and is rapidly expanding toward 10,000+ GPUs, creating unique engineering challenges around performance, efficiency, and hardware utilization.
Rather than joining a large, established ML organization, you'll have significant ownership over architecture, technical direction, and platform decisions from day one.
What You'll Do
- Design and build large-scale serving and inference infrastructure for state-of-the-art ML models.
- Optimize model performance, latency, throughput, and GPU utilization in production environments.
- Deploy and scale models across thousands of GPUs while improving efficiency and reliability.
- Develop model compression, quantization, distillation, and other optimization techniques to maximize hardware performance.
- Work on taking cutting-edge models from research into production by bringing them closer to the hardware.
- Partner closely with researchers and engineers to accelerate experimentation and deployment.
- Help define the architecture and roadmap for a rapidly growing ML platform.
What You Bring
- 8+ years of software engineering, systems engineering, or machine learning infrastructure experience.
- Deep experience with model serving, inference, and performance optimization at scale.
- Strong understanding of model compression, quantization, kernel optimization, and GPU acceleration.
- Experience deploying large-scale ML systems across distributed GPU environments.
- Strong knowledge of CUDA, GPU programming, and hardware-aware optimization.
- Expert-level programming skills in C++ and Python.
- Experience with modern ML frameworks such as PyTorch, JAX, or TensorFlow.
- BS, MS, or PhD in Computer Science, Engineering, Mathematics, or a related field.
Why Consider It?
- Ground-floor opportunity to build a new ML platform inside one of the world's leading trading firms.
- Work alongside engineers and researchers from top AI labs and elite quantitative trading firms.
- Ownership over critical infrastructure supporting next-generation ML workloads.
- Massive scale, with GPU infrastructure growing from the low thousands to 10s of thousands
- Solve challenging problems at the intersection of distributed systems, machine learning, and high-performance computing.
FAQs
Congratulations, we understand that taking the time to apply is a big step. When you apply, your details go directly to the consultant who is sourcing talent. Due to demand, we may not get back to all applicants that have applied. However, we always keep your resume and details on file so when we see similar roles or see skillsets that drive growth in organizations, we will always reach out to discuss opportunities.
Yes. Even if this role isn’t a perfect match, applying allows us to understand your expertise and ambitions, ensuring you're on our radar for the right opportunity when it arises.
We also work in several ways, firstly we advertise our roles available on our site, however, often due to confidentiality we may not post all. We also work with clients who are more focused on skills and understanding what is required to future-proof their business.
That's why we recommend registering your resume so you can be considered for roles that have yet to be created.
Yes, we help with resume and interview preparation. From customized support on how to optimize your resume to interview preparation and compensation negotiations, we advocate for you throughout your next career move.
