Esta vacante solo está disponible en inglés por ahora.

← Volver a Empleos
S

Production Manager / Site Reliability Engineer

Spinomenal·Israel·en
PresencialTiempo completoDevOps & SREiGaming

A digital gaming company is seeking a hands-on Production Manager / SRE to own release engineering, platform stability, incident response, and operational automation in a high-throughput, low-latency environment.

Responsabilidades

  • Validate production configurations and run automated health checks.
  • Own deployment recovery and rollback runbooks.
  • Identify operational toil and automate it with scripting and Infrastructure as Code.
  • Maintain distributed tracing, dashboards, alerting thresholds, and log aggregation.
  • Escalate and troubleshoot complex application, network, database, and infrastructure failures.
  • Lead post-mortems and implement structural fixes to prevent recurring incidents.
  • Govern infrastructure and configuration changes across environments with DevOps partners.
  • Promote SRE practices, operational readiness standards, and resilient architecture across engineering, QA, architecture, DevOps, and production teams.

Requisitos

  • At least 2 years of experience in SRE, production operations, or release engineering for high-transaction web applications.
  • Strong experience building and maintaining Jenkins pipelines, including Pipeline-as-Code and Jenkinsfiles.
  • Hands-on experience managing and troubleshooting AWS infrastructure, including EC2, ECS, S3, Lambda, IAM, and VPC routing.
  • Experience with Terraform and Infrastructure as Code.
  • Proficiency in Python or Bash for automation scripts, system utilities, and internal tools.
  • Advanced experience with observability platforms such as New Relic, Prometheus, and Grafana.
  • Strong SQL skills for log analysis and database troubleshooting.
  • Expertise in web debugging, including JSON payloads, API contracts, HTTP status anomalies, DNS, CDN and Redis caching, and proxies.
  • Understanding of SLIs, SLOs, error budgets, and distributed microservices debugging.
  • Ability to participate in on-call rotations and communicate effectively during critical outages.

Se valora

  • Experience with Ansible or Helm.
  • Experience with Prometheus and Grafana.
  • Experience supporting distributed microservices architectures.

Compatibilidad

Más oportunidades

Vacantes similares

Nuevas vacantes en DevOps & SRE.