Esta vacante solo está disponible en inglés por ahora.

← Volver a Empleos
N

Engineering Manager, Performance Research and Analysis

NVIDIA·Israel·
PresencialTiempo completoEngineering ManagementElectronics Manufacturing

Lead end-to-end performance strategy, research, characterization, testing, and optimization for large-scale AI GPU clusters, networking systems, and data center solutions.

Responsabilidades

  • Drive performance strategy, characterization, test plans, and optimization for AI GPU clusters supporting distributed training and inference.
  • Evaluate and optimize RDMA, networking protocols, collective communication, congestion control, and load-balancing algorithms.
  • Lead performance research for DPUs and storage technologies supporting AI inference deployments.
  • Develop performance observability strategies, telemetry pipelines, dashboards, and automated analytics across NICs, switches, GPUs, and NVLink.
  • Perform root-cause analysis on complex multi-node performance bottlenecks and coordinate mitigation across hardware, firmware, and software teams.

Requisitos

  • Bachelor’s or master’s degree in computer science, computer engineering, software engineering, or equivalent technical experience.
  • At least 8 years of overall experience with deep expertise in high-performance networking, RDMA, and systems-level performance.
  • At least 3 years managing engineering, technical performance, or research and development teams.
  • Hands-on experience analyzing and optimizing collective communication such as NCCL or MPI for distributed AI workloads.
  • Experience analyzing network traffic patterns for large-scale training and inference workloads.
  • Experience designing, deploying, and customizing Grafana dashboards for cluster monitoring, alerting, and data visualization.
  • Strong cross-team leadership, analytical, and communication skills.

Se valora

  • Experience optimizing NCCL, RDMA/RoCEv2, and custom collective algorithms for multi-thousand-GPU deployments.
  • Experience tuning adaptive routing, PFC/ECN congestion control, and packet-spraying technologies.
  • Experience building autonomous performance tools, AI-assisted root-cause analysis agents, or automated regression frameworks.
  • Experience developing Grafana plugins, complex dashboard panels, or alert-management workflows with PromQL or LogQL for hyperscale or HPC environments.

Compatibilidad

Más oportunidades

Vacantes similares

Nuevas vacantes en Engineering Management.

Engineering Manager, AI Cluster Performance and Networking | CVZilla