Senior Performance Engineer – LLM Inference Frameworks
Join an AI infrastructure team developing and optimizing high-performance inference frameworks for large language models running on GPUs.
Обязанности
- Design, implement, and optimize GPU-based LLM inference pipelines
- Profile and tune model execution across schedulers, kernels, and runtime components
- Develop memory management strategies to improve bandwidth use and cache efficiency
- Implement techniques including speculative decoding, context caching, and FP8/INT4 quantization
- Build and maintain benchmarking and testing systems for latency, utilization, and efficiency
Требования
- Bachelor’s, master’s, or higher degree in computer engineering, computer science, applied mathematics, or a related computing field, or equivalent experience
- At least 5 years of relevant software development experience
- Excellent Python, software design, and software engineering skills
- Experience with deep learning frameworks such as PyTorch and Hugging Face
- Experience profiling and debugging Python runtimes, PyTorch internals, and GPU utilization
- Awareness of current LLM architectures and inference techniques
- Fluent written and spoken English
Будет плюсом
- Contributions to TensorRT-LLM, vLLM, SGLang, or similar inference frameworks
- Expertise in performance modeling, memory optimization, distributed model execution, or GPU execution workflows
- Experience with NVIDIA Nsight Systems, PyTorch Profiler, or custom benchmarking tools
- Strong understanding of compute-versus-memory, scheduling-versus-batching, and latency-versus-throughput trade-offs
Условия и преимущества
- Competitive salary
- Comprehensive benefits package