SRE Team Lead
Lead a site reliability engineering team responsible for reliability, observability, automation, and private cloud infrastructure. Own the engineering roadmap and collaborate across development, operations, and platform teams.
Обязанности
- Lead, mentor, and manage a team of SREs while providing technical direction and career development
- Own and prioritize the roadmap for reliability, observability, and automation
- Conduct one-to-ones, performance reviews, and hiring activities
- Promote operational excellence, blameless postmortems, and continuous improvement
- Lead escalations and post-incident reviews for complex reliability issues
- Design and maintain infrastructure automation tools and workflows
- Manage Jenkins and Ansible pipelines using Python and Bash
- Drive infrastructure-as-code adoption
- Improve the scalability, performance, and fault tolerance of critical systems
- Build monitoring and alerting with Grafana, ELK, Zabbix, and custom scripts
- Define and track SLIs, SLOs, and error budgets
- Partner with development teams to embed observability in the software development lifecycle
- Support monitoring and infrastructure integration for MongoDB and PostgreSQL
- Maintain documentation and promote knowledge sharing
Требования
- 5–8+ years of experience in SRE, DevOps, or infrastructure automation
- 3–4+ years of people management or team leadership experience in SRE, DevOps, or infrastructure engineering
- Experience building, coaching, and retaining high-performing engineering teams
- Experience owning engineering roadmaps and leading cross-functional reliability initiatives
- Strong Python and Bash scripting skills
- Hands-on experience with Ansible and infrastructure automation
- Experience with ELK, Grafana, and Zabbix
- Understanding of CI/CD pipelines and Jenkins
- Production experience with MongoDB and PostgreSQL monitoring
- Comfort working in Linux environments
- Strong problem-solving, written communication, and verbal communication skills
Будет плюсом
- Experience with AWS, GCP, or Azure
- Exposure to Docker and Kubernetes
- Familiarity with Terraform
- Experience introducing SLOs, error budgets, or chaos engineering
- Experience using AI tools to improve processes
Условия и преимущества
- Senior leadership role with influence over reliability engineering culture and roadmap
- Collaborative, high-trust environment supporting autonomy, ownership, and learning
- Competitive compensation and benefits
- Professional development support
- Opportunity to shape engineering practices across the organization