DevOps Team Lead
Lead and mentor DevOps engineers while owning the design, deployment, and continuous improvement of secure, scalable, globally distributed cloud infrastructure.
Responsibilities
- Lead and mentor DevOps engineers, providing technical guidance and establishing engineering practices.
- Collaborate with engineering, security, and data teams on scalable and reliable solutions.
- Design, architect, deploy, and maintain AWS and Azure cloud infrastructure.
- Own and improve CI/CD pipelines for automated deployment, testing, and scaling across environments.
- Implement and manage monitoring, logging, and alerting to support reliability and performance.
- Troubleshoot complex infrastructure issues across production, staging, and development environments.
- Promote operational excellence, security best practices, and cloud cost optimization.
- Make technical decisions and balance hands-on delivery with strategic planning.
Requirements
- At least 4 years of experience in DevOps or a related engineering role, including strong AWS and cloud-native experience.
- At least 3 years of people-management and technical leadership experience.
- Experience designing and operating high-scale production systems for reliability, performance, and scalability.
- Strong expertise in cloud technologies, application security, and secure, resilient, cost-efficient architectures.
- Hands-on experience with Docker, Kubernetes, and Helm.
- Experience with Terraform, automation, and CI/CD tools such as Jenkins.
- Experience with monitoring and logging tools such as Prometheus, Grafana, and Loki.
- Proficiency in scripting or programming, including Node.js or Bash.
- Experience managing big-data infrastructure for data-intensive applications.
- Strong communication, collaboration, documentation, ownership, and problem-solving skills.
Nice to have
- Experience with Cassandra, Elasticsearch, ClickHouse, and PostgreSQL high availability.
- Kafka cluster administration, including cross-region replication and high-availability configurations.
- Experience with production MLOps infrastructure, including model deployment, monitoring, and scaling.