Senior DevOps Tech Lead
Lead the technical direction and operation of secure, scalable infrastructure for data and AI platforms. Own architecture decisions, mentor engineers, and drive high-impact infrastructure initiatives across multiple teams and regions.
Responsibilities
- Own architecture and technical roadmaps for critical infrastructure domains
- Manage multi-region Kubernetes clusters, streaming pipelines, and data orchestration platforms
- Build and operate AI gateways, ML inference infrastructure, and AI enablement tooling
- Design autonomous agents and intelligent automation to reduce infrastructure toil
- Lead multi-team infrastructure initiatives from design through rollout
- Mentor DevOps and infrastructure engineers and raise engineering standards
- Build and operate CI/CD and GitOps pipelines using GitHub Actions, ArgoCD, Helm, and Terraform
- Protect sensitive data through access control, governance, and secure infrastructure design
- Evolve data infrastructure using modern data platform technologies
- Improve internal platforms and enable developer self-service
- Provide monitoring, alerting, debugging, and incident response for critical data and AI systems
Requirements
- 6–8+ years of experience in DevOps, infrastructure, or platform engineering
- Production-scale Kubernetes expertise, including cluster management, networking, scaling, and troubleshooting
- Strong Infrastructure as Code and GitOps experience with Terraform, CDKTF, Helm, ArgoCD, or similar tools
- Experience designing, maintaining, and optimizing CI/CD pipelines for multiple teams
- Deep cloud infrastructure experience, preferably AWS, including networking, IAM, security, and cost optimization
- Strong knowledge of application security, access control, and data protection
- Experience with data infrastructure such as Kafka, Airflow, EMR, Apache Iceberg, Snowflake, or ClickHouse
- Fluency in Linux, scripting, and at least one programming language such as Python, TypeScript, or Go
- Ability to design system architecture at scale and make complex technical trade-offs across teams
- Experience owning production reliability, incident leadership, postmortems, and systemic improvements
- Strong communication and cross-functional collaboration skills
- Understanding of software products and experience working with autonomy in ambiguous environments
Nice to have
- AI/ML infrastructure experience, including model serving, LLM deployment, GPU workloads, or AI observability
- Experience building AI agents and infrastructure automation with tools such as LangChain, LangGraph, n8n, or Claude Code
- Experience developing internal platforms and developer self-service tooling