Lead Network Systems Validation Engineer
Lead hands-on validation of advanced networking technologies across large-scale AI clusters, combining technical ownership, automation, debugging, and engineering mentorship.
תחומי אחריות
- Review requirements and define validation methodologies for networking technologies in large-scale AI clusters
- Create and execute functional and performance test plans
- Develop and maintain benchmarks, automation tools, scripts, test environments, log collection, and data-analysis workflows
- Investigate complex issues using logs, telemetry, packet captures, and system metrics; drive defects to root cause and resolution
- Inspect C, C++, and Python code to investigate defects, validate fixes, and improve instrumentation and debugging
- Collaborate with hardware and software teams on NCCL, RoCE, RDMA, and related networking components
- Profile AI training and inference workloads to identify scalability and performance limitations
- Own the validation roadmap, mentor engineers, document findings, and improve validation processes
דרישות
- Bachelor’s degree in Computer Science, Electrical Engineering, or equivalent experience
- At least 12 years of experience in networking, system validation, or related fields
- Strong experience debugging complex production systems through hypothesis-driven experiments and root-cause analysis
- Ability to read, debug, and reason about C and C++ code
- Strong scripting and automation skills with Python, Bash, or Ansible
- Deep understanding of distributed systems, concurrency, consistency models, fault tolerance, and large-scale performance
- Ability to align technical decisions across teams and communicate trade-offs clearly
- Experience or capability in AI-driven test automation, including scenario generation, LLM-assisted root-cause analysis, and autonomous validation pipelines
יתרון
- Experience with large-scale clusters or distributed systems
- Familiarity with ConnectX, BlueField, or related networking solutions
- Experience with performance analysis, Kubernetes, or cloud environments
- Background in chaos testing, fault injection, or simulation systems
- Rust or Go experience