HPC Operations Engineer
Support and improve large-scale high-performance computing infrastructure, from server configuration through application-level troubleshooting. Work with specialist teams to improve reliability and support chip development.
Responsabilidades
- Troubleshoot support requests in a large-scale HPC environment.
- Improve deployment automation, configuration management, observability, and operational monitoring.
- Maintain accurate operating systems and configurations on compute servers.
- Resolve complex issues across hardware and software layers to maintain system reliability and efficiency.
- Coordinate with specialist teams to resolve issues and improve infrastructure use in chip development.
Requisitos
- Bachelor’s degree in computer science or a similar field, or equivalent experience.
- At least 2 years of experience with CentOS or RHEL Linux distributions.
- Understanding of container technologies such as Docker.
- Proficiency in Python and Unix scripting languages such as Bash.
- Experience with cluster configuration management tools such as Ansible.
- Strong troubleshooting and problem-solving skills.
Se valora
- Knowledge of Linux technologies such as NFS, automounter, LDAP, DNS, and TCP/IP.
- Experience administering job schedulers such as IBM Spectrum LSF or SLURM and operating large-scale compute infrastructure.
- Knowledge of FlexLM license management.
- Perl experience maintaining legacy automation scripts.
- Familiarity with high-speed networking such as InfiniBand, RDMA, or RoCE, and distributed storage such as Lustre or GPFS.