Site Reliability Engineer
A global technology organization is seeking a hands-on Site Reliability Engineer to support large-scale, mission-critical production environments. The role focuses on cloud infrastructure, reliability, scalability, automation, and production excellence.
תחומי אחריות
- Own and improve the reliability of production systems
- Automate operational processes and infrastructure
- Build scalable cloud-native infrastructure
- Monitor production environments and improve observability
- Troubleshoot production issues and support incident response
- Conduct root-cause analysis and drive production excellence
דרישות
- Experience in SRE, production engineering, or DevOps
- Hands-on experience with public cloud infrastructure, including AWS, GCP, or Azure
- Production experience with Kubernetes
- Experience with Terraform and infrastructure as code
- Python or similar automation experience
- Experience with CI/CD, observability, and monitoring
- Production troubleshooting, incident response, and root-cause analysis
יתרון
- GCP experience