AI Performance Engineer
Seeking an AI Performance Engineer to optimize LLM inference performance across the full stack, including GPU kernel development, runtime optimization, and distributed serving. Requires strong systems programming, profiling, and performance engineering skills.
Responsibilities
- Optimize LLM inference performance across kernels, runtimes, model execution, networking, scheduling, and distributed serving.
- Analyze and improve GPU utilization, memory bandwidth, latency, throughput, and cost efficiency.
- Develop, tune, or integrate high-performance GPU kernels using CUDA, Triton, CUTLASS, or similar frameworks.
- Improve single-node inference performance through runtime optimization, memory management, batching, quantization, parallelism, and profiling.
- Optimize distributed inference across clusters, including tensor parallelism, pipeline parallelism, expert parallelism, communication patterns, load balancing, and scheduling.
- Build tools, benchmarks, and profiling workflows to identify bottlenecks and measure performance improvements.
- Collaborate with ML, infrastructure, and product teams to translate performance improvements into production impact.
- Stay current with advances in LLM serving, GPU architectures, compiler/runtime systems, and AI infrastructure.
Requirements
- Strong experience in at least one: performance engineering/low-level software optimization, GPU kernel development/optimization, or AI/ML systems optimization.
- Strong programming skills in C++, CUDA, Python, or similar systems-oriented languages.
- Experience profiling and optimizing software for latency, throughput, memory usage, or hardware utilization.
- Familiarity with GPU architectures and performance characteristics.
- Experience with one or more: CUDA, Triton, CUTLASS, ROCm, NCCL, TensorRT, XLA, TVM, vLLM, TensorRT-LLM, PyTorch.
- Understanding of LLM inference techniques: batching, KV cache management, quantization, speculative decoding, parallelism strategies, and distributed serving.
Nice to have
- Experience optimizing transformer-based models or production LLM inference systems.
- Experience with multi-GPU or multi-node inference.
- Familiarity with compiler-level optimization, graph optimization, or model runtime internals.
- Experience with observability, benchmarking, and performance regression testing.
- Contributions to open-source AI infrastructure, GPU computing, or systems performance projects.
- Experience with large-scale distributed systems, high-performance computing, networking, or cluster scheduling.