AI Research Engineer
Implement and evaluate AI model components, with a focus on efficient models and inference. Run reproducible experiments, measure system performance, and work with research and engineering teams to validate designs.
Responsibilities
- Analyze research papers and source code, test assumptions, and turn uncertain claims into hypotheses.
- Implement and modify model components, training loops, and inference paths in Python and PyTorch; verify correctness before optimization.
- Prepare data and environments, run and monitor GPU experiments, recover failed jobs, and preserve reproducible configurations and results.
- Design matched baselines and ablations; evaluate quality and failure modes across model sizes, datasets, and context lengths.
- Measure memory use, data movement, latency, and throughput, and assess whether theoretical improvements translate into system gains.
- Diagnose issues across mathematics, model architecture, data, and implementation; collaborate with research, compiler, runtime, and hardware engineers.
Requirements
- 3–5 years of hands-on AI/ML engineering or applied research experience.
- Strong foundations in linear algebra, probability, optimization, and numerical methods.
- Understanding of Transformer internals, attention, autoregressive inference, and familiarity with recurrent, state-space, or hybrid sequence models.
- Practical depth in at least two areas of efficient AI, such as long-context inference, KV-cache compression, sparse or linear attention, model compression, quantization, conditional computation, or inference optimization.
- Strong Python and PyTorch skills, including model implementation, training and inference code, testing, and debugging.
- Experience running GPU experiments in Linux, profiling workloads, using version control, and managing reproducible results.
- Sound experimental judgment, including the ability to assess generalization, identify confounded comparisons, and explain negative results.
- Ability to work from incomplete specifications, choose useful experiments, and communicate findings clearly.
Nice to have
- Experience with C++, CUDA, Triton, custom GPU kernels, distributed training, compilers, inference runtimes, or hardware-aware algorithm design.
- Experience with model distillation, architecture conversion, fine-tuning, pretraining, scaling prototypes, or research implementations used by others.