Senior AI Inference Engineer
Develop and optimize high-performance AI inference systems across mobile, edge, and distributed hardware environments. The role combines research, low-level engineering, benchmarking, and production integration in a remote-first setting.
Responsibilities
- Design and deploy model-serving architectures optimized for latency, throughput, and memory use.
- Develop inference pipelines for mobile, edge, and other hardware environments.
- Set performance targets and benchmark inference systems in simulated and production settings.
- Create datasets and simulation scenarios for representative model evaluation.
- Identify computational and memory bottlenecks and implement system-level optimizations.
- Develop GPU kernels and compute shaders for mobile hardware, including in MSL.
- Apply inference optimization methods such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.
- Design distributed inference systems for large-scale GPU workloads.
- Integrate optimized inference frameworks into production and edge applications with research and engineering teams.
- Document experimental results, monitor production performance, and refine optimization strategies.
Requirements
- Degree in computer science or a related technical field.
- Expertise in Metal Shading Language (MSL), including writing custom compute shaders from scratch.
- Experience optimizing low-level kernels and inference on mobile or other resource-constrained devices.
- Evidence of measurable improvements to inference latency, throughput, and memory use.
- Strong understanding of model-serving architectures, inference engines, and high-performance AI deployment.
- Experience building and deploying end-to-end inference pipelines on constrained hardware.
- Experience designing inference benchmarking and evaluation frameworks.
- Knowledge of distributed inference methods, including tensor, pipeline, and expert parallelism.
- Understanding of diffusion models and Vision Transformers.
- Familiarity with pruning, quantization, Flash Attention, KV cache optimization, and speculative decoding.
- Strong English communication skills for collaboration with distributed technical teams.
Nice to have
- PhD in NLP, machine learning, or a related discipline.
- AI research track record, including publications at leading conferences.
Benefits
- Remote-first work with an international team.
- Work on advanced AI research and performance-critical infrastructure.
- Contribute to mobile, edge, and distributed inference systems and modern model architectures.
- Collaborate with research and engineering teams on practical AI systems challenges.