Esta vacante solo está disponible en inglés por ahora.

← Volver a Empleos
N

HPC Operations Engineer

NVIDIA·Israel·en
Sin especificarTiempo completoDevOps & SRESemiconductors

Support and improve large-scale high-performance computing infrastructure, from server configuration through application-level troubleshooting. Work with specialist teams to improve reliability and support chip development.

Responsabilidades

  • Troubleshoot support requests in a large-scale HPC environment.
  • Improve deployment automation, configuration management, observability, and operational monitoring.
  • Maintain accurate operating systems and configurations on compute servers.
  • Resolve complex issues across hardware and software layers to maintain system reliability and efficiency.
  • Coordinate with specialist teams to resolve issues and improve infrastructure use in chip development.

Requisitos

  • Bachelor’s degree in computer science or a similar field, or equivalent experience.
  • At least 2 years of experience with CentOS or RHEL Linux distributions.
  • Understanding of container technologies such as Docker.
  • Proficiency in Python and Unix scripting languages such as Bash.
  • Experience with cluster configuration management tools such as Ansible.
  • Strong troubleshooting and problem-solving skills.

Se valora

  • Knowledge of Linux technologies such as NFS, automounter, LDAP, DNS, and TCP/IP.
  • Experience administering job schedulers such as IBM Spectrum LSF or SLURM and operating large-scale compute infrastructure.
  • Knowledge of FlexLM license management.
  • Perl experience maintaining legacy automation scripts.
  • Familiarity with high-speed networking such as InfiniBand, RDMA, or RoCE, and distributed storage such as Lustre or GPFS.

Compatibilidad

Más oportunidades

Vacantes similares

Nuevas vacantes en DevOps & SRE.