NOC Team Lead
Lead global 24/7 NOC operations while advancing monitoring, automation, incident management, and SRE practices for reliable production systems.
Обязанности
- Direct global 24/7 NOC teams and oversee incident response, service uptime, and operational excellence.
- Implement monitoring, alerting, and observability using tools such as Prometheus, Grafana, and Datadog.
- Automate repetitive operational tasks and reduce manual troubleshooting and human error.
- Oversee critical incident management, stakeholder communication, root-cause analysis, and post-incident reviews.
- Coach and develop NOC engineers and SREs while fostering continuous improvement.
- Collaborate with engineering, development, and IT teams to meet performance and SLA requirements.
Требования
- At least 3 years of experience managing technical teams in NOC, SRE, or infrastructure environments.
- Experience with AWS, Azure, or GCP; Kubernetes; Linux or Unix; and Python, Bash, or Golang scripting.
- Experience with monitoring platforms such as Datadog and CI/CD tools such as Jenkins or GitLab.
- Strong communication, leadership, and crisis management skills.