Engineering Manager, Network Systems Validation
Lead a team validating advanced networking solutions across large-scale AI cluster environments. Own the technical roadmap, execution, automation strategy, and engineering quality for network system validation.
Обязанности
- Lead, mentor, and coach software development and system validation engineers
- Define validation methodologies and comprehensive functional and performance test plans for networking technologies
- Develop and maintain benchmarks, automation tools, test scripts, environment setup, log collection, and data analysis systems
- Lead investigations of complex networking issues using logs, telemetry, packet captures, and system metrics
- Triage issues across hardware and software layers and drive them to root cause and resolution
- Collaborate with hardware and software teams to debug networking technologies, including NCCL, RoCE, and RDMA
- Profile AI training and inference workloads and analyze scalability and performance limitations
- Own the team’s technical roadmap, execution, validation processes, and engineering excellence
- Document technical findings and communicate validation results
- Improve automation environments, validation methods, software quality, and engineering processes
Требования
- Bachelor’s degree in Computer Science, Electrical Engineering, or equivalent experience
- At least 8 years of experience in networking, system validation, or related fields
- At least 3 years leading a software or system development team
- Experience debugging complex production systems through hypothesis formation, experimentation, and root-cause analysis
- Strong scripting and automation skills using Python, Bash, and/or Ansible
- Ability to read, debug, and reason about C or C++ code
- Ability to align technical stakeholders, communicate trade-offs, and make architectural decisions
- Experience or knowledge of AI-driven test automation, including scenario generation, LLM-assisted root-cause analysis, or autonomous validation pipelines
Будет плюсом
- Experience with large-scale clusters or distributed systems
- Familiarity with NVIDIA networking solutions such as ConnectX, BlueField, or SpecX
- Experience with performance analysis, Kubernetes, or cloud environments
- Background in chaos testing, fault injection, or simulation systems
- Knowledge of Rust or Go