Senior HPC and AI Infrastructure Operations Engineer
Join an infrastructure team building and operating large-scale HPC and AI clusters. The role supports accelerated computing platforms, research and development, and new infrastructure solutions.
Обязанности
- Deploy, manage, and maintain large-scale HPC and AI clusters
- Manage Linux workload scheduling and orchestration tools
- Support and maintain continuous integration and delivery pipelines
- Troubleshoot systems from bare metal and operating-system layers through software stacks and applications
- Collaborate with HPC, operating-system, GPU computing, and systems specialists to architect and bring up large-scale performance platforms
- Support research and development activities and evaluate future infrastructure improvements
- Work with researchers, developers, and customers to improve workflows and develop differentiated solutions
Требования
- Bachelor’s degree in computer science, engineering, or a related field, or equivalent experience
- At least 5 years of professional experience
- Knowledge of HPC and AI technologies spanning CPUs, GPUs, high-speed interconnects, and supporting software
- Experience with workload scheduling and orchestration tools such as Slurm and Kubernetes
- Strong knowledge of Windows and Linux systems, including Red Hat, CentOS, and Ubuntu
- Knowledge of networking, sockets, firewalls, iptables, Wireshark, ACLs, operating-system security, TCP, DHCP, and DNS
- Experience with storage solutions such as Lustre, GPFS, ZFS, and XFS
- Python programming and Bash scripting experience
- Experience with Jenkins, Ansible, GitOps, or comparable automation and configuration-management tools
- Knowledge of InfiniBand and Ethernet networking protocols
- Experience with virtualized systems such as VMware, Hyper-V, or KVM
- Familiarity with AWS, Azure, or Google Cloud
Будет плюсом
- Knowledge of CPU or GPU architecture
- Kubernetes and containerized microservices experience
- Experience with GPU-focused hardware and software such as DGX and CUDA
- Experience with RDMA fabrics, including InfiniBand or RoCE
Условия и преимущества
- Competitive salary
- Extensive benefits package
- Flexible work environment
- Inclusive and supportive workplace