Senior Software Engineer, LLM Inference
Develop and optimize production-grade LLM inference software, GPU kernels, and distributed systems across data center and edge hardware platforms.
Responsibilities
- Implement and optimize inference algorithms for LLM and multimodal architectures
- Profile inference pipelines and correlate simulation results with real hardware performance
- Write and tune CUDA and Triton GPU kernels for inference operators
- Solve distributed inference challenges involving parallelism, communication, and multi-node deployment
- Build production software in open-source inference libraries and runtimes
- Own optimization features from scoping through delivery while collaborating across research, product, and engineering teams
Requirements
- B.Sc., M.Sc., or equivalent experience in Computer Science, Computer Engineering, or a related technical field
- At least 5 years of hands-on software engineering experience in performance-critical systems
- Understanding of deep learning architectures, including Transformers, SSMs, and mixture-of-experts models
- Experience with GPU programming, memory hierarchy, networking, or distributed computing
- Experience optimizing deep learning workloads on NVIDIA GPUs using profiling and tracing tools
- Strong software engineering fundamentals, including clean design, extensibility, and testability
Nice to have
- Contributions to open-source inference runtimes such as vLLM, SGLang, FlashInfer, or Dynamo
- Experience with LLM quantization, mixed precision, and numerical tradeoffs
- Distributed inference at scale, including tensor, pipeline, or expert parallelism
- Knowledge of emerging LLM architectures and state-space mechanisms
- Experience with performance modeling and simulation-to-hardware correlation