Senior Systems Software Engineer – Network and Collectives
Join an engineering team developing the scale-out communication substrate for AI accelerator systems. You will build topology-aware collective communication software across multi-node pods and optimize performance across the underlying interconnect fabric.
תחומי אחריות
- Design and implement collective algorithms tuned to interconnect topology and bandwidth and latency characteristics.
- Build a topology-aware transport layer across tray, rack, pod, and cluster interconnects.
- Optimize end-to-end collective performance across multi-node pods by profiling bottlenecks and improving fabric utilization.
- Collaborate with hardware engineers on interconnect features and with runtime engineers on partitioning, overlap, and scheduling.
- Own correctness and numerical determinism for reductions at scale.
דרישות
- At least 5 years of experience in AI, systems, or HPC software development.
- Strong C++ programming skills and experience with concurrency and lock-free design.
- Hands-on experience with collective libraries such as NCCL or MPI, or with scale-out communication systems.
- Understanding of tensor, pipeline, and expert parallelism and how they map to collectives and hardware fabrics.
- Experience with profiling, bandwidth and latency tuning, and roofline-based performance analysis.
יתרון
- Experience with topology or placement algorithms, congestion control, or multi-tenant fabrics.
- Experience with large-scale distributed training or inference.
- Experience with inference serving stacks such as vLLM, SGLang, TensorRT-LLM, or DeepSeek.
- Experience with RDMA, InfiniBand, RoCE, GPUDirect-style transfers, or comparable fabric technologies.