Senior Site Reliability Engineer
A cybersecurity company is seeking a Senior Site Reliability Engineer to design reliable infrastructure, automate operations, and improve observability across complex customer environments.
Обязанности
- Design and develop infrastructure and reliability solutions from architecture through production deployment.
- Automate deployment, configuration, provisioning, and product and infrastructure lifecycle management.
- Design and build observability solutions for system health, performance, reliability, and customer experience.
- Develop metrics, logs, events, dashboards, and alerting to detect failures and reliability issues.
- Build internal tools for debugging, diagnostics, troubleshooting, and operational visibility across application, container, system, and network layers.
- Investigate complex system and networking issues, perform root cause analysis, and drive problems to resolution.
- Collaborate with research and development, deployment, support, and other engineering teams to improve reliability, scalability, observability, and operational efficiency.
Требования
- At least 3 years of hands-on experience in SRE, DevOps, or infrastructure engineering in production environments.
- Strong Python software engineering skills, including production-quality automation, plus Bash scripting experience.
- Strong Linux expertise covering troubleshooting, networking, processes, configuration, and kernel behavior.
- Strong networking knowledge, including TCP/IP, DNS, routing, NAT, VPNs, firewalls, and proxies.
- Hands-on Kubernetes experience with cluster troubleshooting, networking, deployments, and Helm chart development and maintenance.
- Experience with observability systems covering metrics, logs, dashboards, alerting, and end-to-end troubleshooting.
- Experience with Prometheus, Grafana, Elastic, Fluentd, Vector, or comparable technologies.
- Experience with Ansible or similar infrastructure automation tools.
- Strong debugging and root cause analysis skills across multiple infrastructure layers.
- Ability to independently investigate and resolve complex technical problems.
- Strong communication and collaboration skills across engineering, deployment, and support teams.
Будет плюсом
- Experience with Terraform, Terragrunt, or similar infrastructure-as-code technologies.