← Back to jobs
G

Staff Site Reliability Engineer

Grubhub·Tel Aviv, Israel·en
Not specifiedFull-timeDevOps & SREFood Delivery

Own reliability and technical direction for a high-traffic platform, covering AWS infrastructure, Kubernetes, multi-region resilience, observability, and incident management.

Responsibilities

  • Architect resilient, self-healing production systems and co-own critical service design.
  • Own multi-region resilience, including failover readiness, runbooks, drills, recovery targets, and data replication.
  • Design, deploy, and maintain AWS infrastructure as code using Terraform or similar tools.
  • Own the EKS platform lifecycle, including upgrades, ingress, and autoscaling.
  • Manage observability pipelines for logs, metrics, tracing, and alerts, including signal quality and cost.
  • Use SLOs and telemetry to improve reliability and prevent incidents.
  • Set scaling and capacity strategy for seasonal traffic peaks.
  • Manage cloud costs through right-sizing, reservations, and savings tracking.
  • Build and maintain CI/CD pipelines and deployment tooling.
  • Operate within PCI-scoped environments and their access controls.
  • Lead incident response, postmortems, failure analysis, service reviews, and architecture reviews.

Requirements

  • Typically 8+ years in SRE, DevOps, or infrastructure engineering; demonstrated staff-level ownership and impact are emphasized over tenure.
  • Experience owning a production platform end to end and setting technical direction across teams.
  • Deep Infrastructure as Code experience, especially Terraform and multi-environment rollout.
  • Operator-level Kubernetes experience, including EKS, Helm, controllers, ingress, and autoscaling.
  • Experience with multi-region AWS architecture, failover, and data replication.
  • Experience building CI/CD pipelines with tools such as Jenkins or GitHub Actions.
  • Software engineering experience in Python, Go, or a similar object-oriented language.
  • Experience with MySQL, MongoDB or Atlas, Redis or ElastiCache, and message brokers such as RabbitMQ, AmazonMQ, or SQS.
  • Knowledge of microservice architecture, application design, distributed monitoring, SLOs, metrics, tracing, and log pipelines.
  • Working knowledge of AWS, Linux, storage, networking, and PCI-scoped or equivalent environments.
  • Strong technical writing, documentation, and communication skills; able to build consensus across teams in multiple time zones.
  • Experience with highly trafficked web services.

Benefits

  • Equity and 401(k)
  • Medical, dental, and vision plans
  • Company-paid short- and long-term disability coverage
  • Paid time off, including flexible time off for exempt employees, vacation for non-exempt employees, and paid sick leave
  • Paid parental leave
  • Discounted meals and company perks

Relevance

More opportunities

Similar jobs

The newest open roles in DevOps & SRE.

Questions, answered

Frequently asked questions