Senior Performance Engineer - LLM Inference Frameworks
Join a team building core inference infrastructure for large language models. You will design and optimize high-performance GPU pipelines, profile execution, and implement cutting-edge techniques like speculative decoding and quantization to maximize throughput and efficiency.
Responsibilities
- Design, implement, and optimize high-performance inference pipelines for large language models on GPUs
- Profile and tune model execution across the stack, from scheduler design to kernel fusions
- Design and experiment with memory management strategies for improved bandwidth and cache efficiency
- Implement techniques such as Speculative Decoding, Context Caching, and FP8/INT4 quantization
- Develop and maintain benchmarking and testing systems to quantify latency, utilization, and efficiency
Requirements
- Bachelor's degree or higher in Computer Engineering, Computer Science, Applied Mathematics, or related field
- 5+ years of relevant software development experience
- Excellent Python programming and software engineering skills
- Experience with deep learning frameworks like PyTorch and HuggingFace
- Experience profiling and debugging at Python runtime, PyTorch internals, and GPU utilization levels
- Awareness of latest LLM architectures and inference techniques
- Proactive and able to work independently
- Excellent written and oral communication skills in English
Nice to have
- Contributions to inference frameworks such as TensorRT-LLM, vLLM, SGLang, or similar
- Expertise in performance modeling, memory optimization, distributed model execution, or GPU workflows
- Hands-on experience with NVIDIA profiling tools (Nsight Systems, PyTorch Profiler, custom harnesses)
- Strong grasp of trade-offs in inference efficiency: compute vs. memory, scheduling vs. batching, latency vs. throughput