המשרה הזו זמינה כרגע באנגלית.

חזרה למשרות
N

Engineering Manager, Performance Research and Analysis

NVIDIA·ישראל·
במקום העבודהמשרה מלאהEngineering ManagementElectronics Manufacturing

Lead end-to-end performance strategy, research, characterization, testing, and optimization for large-scale AI GPU clusters, networking systems, and data center solutions.

תחומי אחריות

  • Drive performance strategy, characterization, test plans, and optimization for AI GPU clusters supporting distributed training and inference.
  • Evaluate and optimize RDMA, networking protocols, collective communication, congestion control, and load-balancing algorithms.
  • Lead performance research for DPUs and storage technologies supporting AI inference deployments.
  • Develop performance observability strategies, telemetry pipelines, dashboards, and automated analytics across NICs, switches, GPUs, and NVLink.
  • Perform root-cause analysis on complex multi-node performance bottlenecks and coordinate mitigation across hardware, firmware, and software teams.

דרישות

  • Bachelor’s or master’s degree in computer science, computer engineering, software engineering, or equivalent technical experience.
  • At least 8 years of overall experience with deep expertise in high-performance networking, RDMA, and systems-level performance.
  • At least 3 years managing engineering, technical performance, or research and development teams.
  • Hands-on experience analyzing and optimizing collective communication such as NCCL or MPI for distributed AI workloads.
  • Experience analyzing network traffic patterns for large-scale training and inference workloads.
  • Experience designing, deploying, and customizing Grafana dashboards for cluster monitoring, alerting, and data visualization.
  • Strong cross-team leadership, analytical, and communication skills.

יתרון

  • Experience optimizing NCCL, RDMA/RoCEv2, and custom collective algorithms for multi-thousand-GPU deployments.
  • Experience tuning adaptive routing, PFC/ECN congestion control, and packet-spraying technologies.
  • Experience building autonomous performance tools, AI-assisted root-cause analysis agents, or automated regression frameworks.
  • Experience developing Grafana plugins, complex dashboard panels, or alert-management workflows with PromQL or LogQL for hyperscale or HPC environments.

רלוונטיות

הזדמנויות נוספות

משרות דומות

המשרות הפתוחות החדשות ביותר בתחום Engineering Management.

צור פרופיל כדי לראות את מידת ההתאמה שלך.

Engineering Manager, AI Cluster Performance and Networking | CVZilla