← Back to jobs
N

Senior HPC and AI Operations Engineer

NVIDIA·Israel·en
HybridFull-timeDevOps & SREElectronics ManufacturingEnterprise Software

Join an infrastructure team building and operating large-scale high-performance computing and AI clusters. This hybrid role supports accelerated computing platforms, research initiatives, and customer-facing infrastructure solutions in Yokneam Illit, Israel.

Responsibilities

  • Deploy, manage, and maintain large-scale HPC and AI clusters
  • Manage Linux workload scheduling and orchestration tools
  • Support and maintain continuous integration and delivery pipelines
  • Troubleshoot infrastructure from bare metal and operating systems through software stacks and applications
  • Collaborate with HPC, operating system, GPU, and systems specialists to architect and bring up large-scale performance platforms
  • Support research and development activities and participate in proofs of concept and proofs of value
  • Work with researchers, developers, and customers to improve workflows and develop differentiated infrastructure solutions

Requirements

  • Bachelor’s degree in computer science, engineering, or a related field, or equivalent experience
  • At least 5 years of relevant experience
  • Knowledge of HPC and AI technologies spanning CPUs, GPUs, high-speed interconnects, and supporting software
  • Experience with workload scheduling and orchestration tools such as Slurm and Kubernetes
  • Strong knowledge of Windows and Linux systems, including Red Hat/CentOS and Ubuntu
  • Knowledge of networking, sockets, firewalls, iptables, Wireshark, ACLs, operating system security, TCP, DHCP, and DNS
  • Experience with storage solutions such as Lustre, GPFS, ZFS, and XFS
  • Python programming and Bash scripting experience
  • Experience with automation and configuration management tools such as Jenkins, Ansible, and GitOps
  • Knowledge of InfiniBand and Ethernet networking protocols
  • Experience with virtual systems such as VMware, Hyper-V, or KVM
  • Familiarity with cloud platforms including AWS, Azure, or Google Cloud

Nice to have

  • Knowledge of CPU or GPU architecture
  • Kubernetes and containerized microservice technologies
  • GPU-focused platforms such as DGX and CUDA
  • RDMA fabrics, including InfiniBand or RoCE

Benefits

  • Competitive salary
  • Extensive benefits package
  • Flexible work environment
  • Diversity and inclusion initiatives

Relevance

More opportunities

Similar jobs

The newest open roles in DevOps & SRE.

Questions, answered

Frequently asked questions