Senior HPC DevOps Engineer
Build and operate large-scale HPC and AI clusters, collaborating with specialists and researchers to improve system design, performance, and workflows.
Обязанности
- Design, implement, and maintain large-scale HPC and AI clusters
- Build monitoring, logging, and alerting systems
- Develop infrastructure-as-code tools for scalable, repeatable deployments
- Develop and maintain CI/CD pipelines
- Automate deployment, configuration management, and operational monitoring
- Develop networking automation
- Troubleshoot systems from bare metal through application level
- Share technical guidance and best practices with internal teams
- Support research and development initiatives, including proofs of concept and value
Требования
- Bachelor’s degree in computer science, engineering, or a related field
- At least 5 years of experience
- Advanced programming and scripting skills, including object-oriented programming
- Familiarity with Jenkins, Ansible, and Puppet or Chef
- Strong understanding of Kubernetes and container-based microservices
- Hands-on experience with event streaming or message queues such as Apache Kafka
- Experience with storage solutions such as Lustre, GPFS, ZFS, and XFS
- Experience with virtualization systems such as VMware, Hyper-V, KVM, or Citrix
- Familiarity with AWS, Azure, or Google Cloud
Будет плюсом
- Networking experience or relevant professional training
- Knowledge of CPU and/or GPU architecture
- Experience with workload scheduling tools such as Slurm