Senior HPC DevOps Engineer
An established computing hardware and accelerated-computing company is seeking a Senior HPC DevOps Engineer to build, operate, and optimize large-scale HPC and AI platforms. The role works with engineering teams, researchers, developers, and customers on cluster architecture, automation, reliability
Responsabilidades
- Design, implement, and maintain large-scale HPC and AI clusters
- Develop monitoring, logging, and alerting systems
- Create infrastructure-as-code tools for scalable, repeatable deployments
- Develop and maintain CI/CD pipelines
- Automate deployment, configuration management, and operational monitoring
- Develop complex network automation
- Troubleshoot issues from bare metal through the application layer
- Define and share operational best practices with internal teams
- Support research and development activities, proofs of concept, and proofs of value
Requisitos
- Bachelor’s degree in Computer Science, Engineering, or a related field
- At least 5 years of professional experience
- Advanced programming and scripting skills with knowledge of object-oriented programming
- Experience with Jenkins, Ansible, and Puppet or Chef
- Strong understanding of Kubernetes and containerized microservices
- Hands-on experience with event-streaming or message-queue technologies such as Apache Kafka
- Experience with storage technologies including Lustre, GPFS, ZFS, and XFS
- Experience with virtualization systems such as VMware, Hyper-V, KVM, or Citrix
- Familiarity with AWS, Azure, or Google Cloud
Se valora
- Professional networking experience or formal networking training
- Knowledge of CPU or GPU architecture
- Experience with Slurm and Kubernetes for workload scheduling and orchestration
Beneficios
- Inclusive workplace
- Reasonable accommodations for applicants and employees