Эта вакансия пока только на английском.

← Назад к вакансиям
S

Production Manager / Site Reliability Engineer

Spinomenal·Израиль·en
В офисеПолная занятостьDevOps & SREiGaming

A digital gaming company is seeking a hands-on Production Manager / SRE to own release engineering, platform stability, incident response, and operational automation in a high-throughput, low-latency environment.

Обязанности

  • Validate production configurations and run automated health checks.
  • Own deployment recovery and rollback runbooks.
  • Identify operational toil and automate it with scripting and Infrastructure as Code.
  • Maintain distributed tracing, dashboards, alerting thresholds, and log aggregation.
  • Escalate and troubleshoot complex application, network, database, and infrastructure failures.
  • Lead post-mortems and implement structural fixes to prevent recurring incidents.
  • Govern infrastructure and configuration changes across environments with DevOps partners.
  • Promote SRE practices, operational readiness standards, and resilient architecture across engineering, QA, architecture, DevOps, and production teams.

Требования

  • At least 2 years of experience in SRE, production operations, or release engineering for high-transaction web applications.
  • Strong experience building and maintaining Jenkins pipelines, including Pipeline-as-Code and Jenkinsfiles.
  • Hands-on experience managing and troubleshooting AWS infrastructure, including EC2, ECS, S3, Lambda, IAM, and VPC routing.
  • Experience with Terraform and Infrastructure as Code.
  • Proficiency in Python or Bash for automation scripts, system utilities, and internal tools.
  • Advanced experience with observability platforms such as New Relic, Prometheus, and Grafana.
  • Strong SQL skills for log analysis and database troubleshooting.
  • Expertise in web debugging, including JSON payloads, API contracts, HTTP status anomalies, DNS, CDN and Redis caching, and proxies.
  • Understanding of SLIs, SLOs, error budgets, and distributed microservices debugging.
  • Ability to participate in on-call rotations and communicate effectively during critical outages.

Будет плюсом

  • Experience with Ansible or Helm.
  • Experience with Prometheus and Grafana.
  • Experience supporting distributed microservices architectures.

Соответствие

Больше возможностей

Похожие вакансии

Новые вакансии в категории «DevOps & SRE».

Создайте профиль, чтобы увидеть оценку соответствия.

Создайте профиль, чтобы увидеть оценку соответствия.

Создайте профиль, чтобы увидеть оценку соответствия.

Создайте профиль, чтобы увидеть оценку соответствия.