ML Systems Engineer
Develop and optimize deep learning systems for high-frequency trading, working on distributed training, model serving, and performance profiling across massive compute clusters.
Responsibilities
- Develop and maintain large-scale deep learning systems in production
- Collaborate with researchers and engineers to run deep learning models on compute clusters
- Optimize performance of deep learning workloads through profiling and tuning
- Orchestrate and scale distributed training across hundreds to thousands of GPUs
- Optimize model serving and inference pipelines using quantization, distillation, and compilation
- Adapt models for strict production serving constraints
Requirements
- B.Sc. with honors in CS, EE, Math, Physics, or related field from a top-tier university
- 5+ years of experience building and deploying large-scale deep learning systems
- Advanced proficiency in PyTorch or TensorFlow
Nice to have
- M.Sc. or Ph.D. in a quantitative field
- Proficiency in Python, C, or C++
- Deep knowledge of PyTorch internals
- Experience with performance profiling and optimization of deep learning workloads
- Experience orchestrating large-scale distributed training
- Experience optimizing model serving (quantization, distillation, compilation, memory optimization)
- Experience training and scaling state-of-the-art vision, language, or diffusion models
- Experience implementing custom CUDA or Triton kernels