Senior Site Reliability Engineer
Manage and improve hybrid, multi-cloud and on-premises infrastructure supporting large-scale AI and HPC workloads. Own reliability, performance, troubleshooting, automation and observability across Linux systems, Kubernetes clusters and high-density accelerator hardware.
Responsibilities
- Build and maintain infrastructure for large-scale AI and HPC workloads across on-premises and cloud environments
- Operate and enhance multi-cloud, multi-cluster scheduling platforms
- Troubleshoot kernel, driver, networking, storage and distributed-system performance issues
- Maintain critical queuing, time-series database and logging services
- Develop integrated automation and infrastructure tooling
- Optimize hardware utilization with ML and IT engineering teams
- Establish best practices for system design, observability and infrastructure as code
Requirements
- 10+ years of hands-on experience in SRE, Linux administration or systems engineering
- Expert knowledge of Linux internals, debugging and performance tuning
- Experience managing Kubernetes at scale, including managed EKS and bare-metal deployments
- Hands-on experience with distributed queuing, observability, logging and database systems
- Proficiency with Terraform, Helm and configuration management
- Strong networking fundamentals and Bash proficiency
Nice to have
- Experience with GPU or accelerator scheduling and AI/ML pipelines
- Experience with multi-cloud architectures and hybrid environments
- Experience with workflow orchestration tools such as Argo Workflows