Senior Performance Engineer
A semiconductor technology company is seeking a Senior Performance Engineer to develop performance analysis systems for large-scale GPU and CPU clusters supporting AI and high-performance computing. The role works across hardware, firmware, networking, and software teams in a fast-moving R&D setting
תחומי אחריות
- Profile, benchmark, and analyze AI and HPC workloads on GPU and CPU clusters
- Analyze high-performance networking and collective communications, including NCCL, RDMA, MPI, and RoCE
- Identify bottlenecks across networking, compute, memory, and system architecture
- Develop and improve performance analysis, benchmarking, and diagnostic tools
- Define performance test plans and expectations for new technologies and platforms
- Collaborate with hardware, firmware, networking, systems, and software teams
- Support telemetry collection and data refinement for performance analysis
- Maintain data quality, reproducibility, and traceability of performance results
דרישות
- B.Sc. or M.Sc. in Computer Science, Computer Engineering, Software Engineering, or equivalent experience
- At least 5 years of experience in performance analysis, systems engineering, or HPC and AI infrastructure
- Expertise in performance analysis methodologies
- Hands-on experience with high-performance networking, including RDMA, MPI, NCCL, or congestion control
- Strong understanding of latency, throughput, and resource-utilization metrics
- Exposure to hardware, firmware, or embedded telemetry environments
- Strong analytical, problem-solving, and communication skills
- Ability to collaborate effectively in cross-functional R&D teams
יתרון
- Knowledge of CUDA, NCCL internals, and congestion control algorithms
- Deep understanding of CPU architectures, GPUs, HCAs, memory, and PCIe
- Experience with NVIDIA GPUs, CUDA, PyTorch, or TensorFlow
- Experience with cloud platforms
- Proficiency in Python; Bash and C/C++ experience
- Strong experience working in Linux environments