Production Manager / Site Reliability Engineer
A digital gaming company is seeking a hands-on Production Manager / SRE to own release engineering, platform stability, incident response, and operational automation in a high-throughput, low-latency environment.
תחומי אחריות
- Validate production configurations and run automated health checks.
- Own deployment recovery and rollback runbooks.
- Identify operational toil and automate it with scripting and Infrastructure as Code.
- Maintain distributed tracing, dashboards, alerting thresholds, and log aggregation.
- Escalate and troubleshoot complex application, network, database, and infrastructure failures.
- Lead post-mortems and implement structural fixes to prevent recurring incidents.
- Govern infrastructure and configuration changes across environments with DevOps partners.
- Promote SRE practices, operational readiness standards, and resilient architecture across engineering, QA, architecture, DevOps, and production teams.
דרישות
- At least 2 years of experience in SRE, production operations, or release engineering for high-transaction web applications.
- Strong experience building and maintaining Jenkins pipelines, including Pipeline-as-Code and Jenkinsfiles.
- Hands-on experience managing and troubleshooting AWS infrastructure, including EC2, ECS, S3, Lambda, IAM, and VPC routing.
- Experience with Terraform and Infrastructure as Code.
- Proficiency in Python or Bash for automation scripts, system utilities, and internal tools.
- Advanced experience with observability platforms such as New Relic, Prometheus, and Grafana.
- Strong SQL skills for log analysis and database troubleshooting.
- Expertise in web debugging, including JSON payloads, API contracts, HTTP status anomalies, DNS, CDN and Redis caching, and proxies.
- Understanding of SLIs, SLOs, error budgets, and distributed microservices debugging.
- Ability to participate in on-call rotations and communicate effectively during critical outages.
יתרון
- Experience with Ansible or Helm.
- Experience with Prometheus and Grafana.
- Experience supporting distributed microservices architectures.