← Back to jobs
N

Senior HPC and AI Infrastructure Operations Engineer

NVIDIA·Israel·en
HybridFull-timeDevOps & SREElectronics Manufacturing

Join an infrastructure team building and operating large-scale HPC and AI clusters. The role supports accelerated computing platforms, research and development, and new infrastructure solutions.

Responsibilities

  • Deploy, manage, and maintain large-scale HPC and AI clusters
  • Manage Linux workload scheduling and orchestration tools
  • Support and maintain continuous integration and delivery pipelines
  • Troubleshoot systems from bare metal and operating-system layers through software stacks and applications
  • Collaborate with HPC, operating-system, GPU computing, and systems specialists to architect and bring up large-scale performance platforms
  • Support research and development activities and evaluate future infrastructure improvements
  • Work with researchers, developers, and customers to improve workflows and develop differentiated solutions

Requirements

  • Bachelor’s degree in computer science, engineering, or a related field, or equivalent experience
  • At least 5 years of professional experience
  • Knowledge of HPC and AI technologies spanning CPUs, GPUs, high-speed interconnects, and supporting software
  • Experience with workload scheduling and orchestration tools such as Slurm and Kubernetes
  • Strong knowledge of Windows and Linux systems, including Red Hat, CentOS, and Ubuntu
  • Knowledge of networking, sockets, firewalls, iptables, Wireshark, ACLs, operating-system security, TCP, DHCP, and DNS
  • Experience with storage solutions such as Lustre, GPFS, ZFS, and XFS
  • Python programming and Bash scripting experience
  • Experience with Jenkins, Ansible, GitOps, or comparable automation and configuration-management tools
  • Knowledge of InfiniBand and Ethernet networking protocols
  • Experience with virtualized systems such as VMware, Hyper-V, or KVM
  • Familiarity with AWS, Azure, or Google Cloud

Nice to have

  • Knowledge of CPU or GPU architecture
  • Kubernetes and containerized microservices experience
  • Experience with GPU-focused hardware and software such as DGX and CUDA
  • Experience with RDMA fabrics, including InfiniBand or RoCE

Benefits

  • Competitive salary
  • Extensive benefits package
  • Flexible work environment
  • Inclusive and supportive workplace

Relevance

More opportunities

Similar jobs

The newest open roles in DevOps & SRE.

Questions, answered

Frequently asked questions