Senior Performance Engineer, LLM Inference Frameworks
Build and optimize high-performance inference infrastructure for large language models running on GPUs. Improve throughput, latency, memory efficiency, and scalability across modern model runtimes.
תחומי אחריות
- Design, implement, and optimize GPU-based LLM inference pipelines
- Profile and tune model execution across schedulers, kernels, and runtime components
- Develop memory management strategies to improve bandwidth utilization and cache efficiency
- Implement speculative decoding, context caching, FP8, INT4 quantization, and related inference optimizations
- Build and maintain benchmarking and testing systems for latency, utilization, and efficiency
דרישות
- Bachelor’s, master’s, or higher degree in computer engineering, computer science, applied mathematics, or a related computing field, or equivalent experience
- At least 5 years of relevant software development experience
- Excellent Python programming, software design, and software engineering skills
- Experience with deep learning frameworks including PyTorch and Hugging Face
- Experience profiling and debugging Python runtimes, PyTorch internals, and GPU utilization
- Awareness of current LLM architectures and inference techniques
- Excellent written and spoken English
יתרון
- Contributions to inference frameworks such as TensorRT-LLM, vLLM, or SGLang
- Expertise in performance modeling, memory optimization, distributed model execution, or GPU execution workflows
- Experience with NVIDIA Nsight Systems, PyTorch Profiler, or custom benchmarking harnesses
- Strong understanding of compute, memory, scheduling, batching, latency, and throughput trade-offs
הטבות
- Competitive salary
- Comprehensive benefits package