Engineering Manager, Performance Research and Analysis
Lead end-to-end performance strategy, research, characterization, testing, and optimization for large-scale AI GPU clusters, networking systems, and data center solutions.
Обязанности
- Drive performance strategy, characterization, test plans, and optimization for AI GPU clusters supporting distributed training and inference.
- Evaluate and optimize RDMA, networking protocols, collective communication, congestion control, and load-balancing algorithms.
- Lead performance research for DPUs and storage technologies supporting AI inference deployments.
- Develop performance observability strategies, telemetry pipelines, dashboards, and automated analytics across NICs, switches, GPUs, and NVLink.
- Perform root-cause analysis on complex multi-node performance bottlenecks and coordinate mitigation across hardware, firmware, and software teams.
Требования
- Bachelor’s or master’s degree in computer science, computer engineering, software engineering, or equivalent technical experience.
- At least 8 years of overall experience with deep expertise in high-performance networking, RDMA, and systems-level performance.
- At least 3 years managing engineering, technical performance, or research and development teams.
- Hands-on experience analyzing and optimizing collective communication such as NCCL or MPI for distributed AI workloads.
- Experience analyzing network traffic patterns for large-scale training and inference workloads.
- Experience designing, deploying, and customizing Grafana dashboards for cluster monitoring, alerting, and data visualization.
- Strong cross-team leadership, analytical, and communication skills.
Будет плюсом
- Experience optimizing NCCL, RDMA/RoCEv2, and custom collective algorithms for multi-thousand-GPU deployments.
- Experience tuning adaptive routing, PFC/ECN congestion control, and packet-spraying technologies.
- Experience building autonomous performance tools, AI-assisted root-cause analysis agents, or automated regression frameworks.
- Experience developing Grafana plugins, complex dashboard panels, or alert-management workflows with PromQL or LogQL for hyperscale or HPC environments.