DevOps Infrastructure Team Lead
Lead and develop a team responsible for DevOps and infrastructure engineering, including scalable production systems, Kubernetes platforms, development and production infrastructure, and self-hosted AI environments.
Responsibilities
- Lead, mentor, and grow a team of DevOps and infrastructure engineers
- Own the architecture, development, and operation of development and production infrastructure
- Manage self-hosted AI environments, including on-premises GPU infrastructure
- Lead Kubernetes infrastructure design, development, deployment, operations, troubleshooting, and lifecycle management
- Define infrastructure standards, best practices, and engineering methodologies
- Own the DevOps toolchain and monitoring and logging platforms
- Partner with engineering teams to enable delivery and resolve platform challenges
- Evaluate and introduce technologies through proofs of concept and technical assessments
- Drive incident resolution, root-cause analysis, and continuous infrastructure improvement
Requirements
- At least 5 years of hands-on DevOps engineering experience
- At least 2 years of experience managing, leading, or mentoring engineers
- Experience designing, managing, and maintaining highly scalable production systems
- Hands-on production experience with Kubernetes
- Strong experience with Terraform, Ansible, and Helm
- Experience with Prometheus, Grafana, and ELK or comparable monitoring, observability, and logging stacks
- Strong programming or scripting skills in Python, Go, Ruby, Java, PowerShell, or Bash
- Experience with GitOps practices and tools such as Argo CD
- Strong technical leadership, problem-solving, and communication skills
- A proactive, independent, hands-on approach and willingness to learn
Nice to have
- Experience managing on-premises GPU infrastructure for AI and machine learning environments
- Strong knowledge of networking fundamentals
- Experience defining technical strategy, architecture, and infrastructure roadmaps
- Experience with highly available, large-scale, or security-sensitive environments