Staff Site Reliability Engineer
Own reliability and technical direction for a high-traffic platform, covering AWS infrastructure, Kubernetes, multi-region resilience, observability, and incident management.
Responsibilities
- Architect resilient, self-healing production systems and co-own critical service design.
- Own multi-region resilience, including failover readiness, runbooks, drills, recovery targets, and data replication.
- Design, deploy, and maintain AWS infrastructure as code using Terraform or similar tools.
- Own the EKS platform lifecycle, including upgrades, ingress, and autoscaling.
- Manage observability pipelines for logs, metrics, tracing, and alerts, including signal quality and cost.
- Use SLOs and telemetry to improve reliability and prevent incidents.
- Set scaling and capacity strategy for seasonal traffic peaks.
- Manage cloud costs through right-sizing, reservations, and savings tracking.
- Build and maintain CI/CD pipelines and deployment tooling.
- Operate within PCI-scoped environments and their access controls.
- Lead incident response, postmortems, failure analysis, service reviews, and architecture reviews.
Requirements
- Typically 8+ years in SRE, DevOps, or infrastructure engineering; demonstrated staff-level ownership and impact are emphasized over tenure.
- Experience owning a production platform end to end and setting technical direction across teams.
- Deep Infrastructure as Code experience, especially Terraform and multi-environment rollout.
- Operator-level Kubernetes experience, including EKS, Helm, controllers, ingress, and autoscaling.
- Experience with multi-region AWS architecture, failover, and data replication.
- Experience building CI/CD pipelines with tools such as Jenkins or GitHub Actions.
- Software engineering experience in Python, Go, or a similar object-oriented language.
- Experience with MySQL, MongoDB or Atlas, Redis or ElastiCache, and message brokers such as RabbitMQ, AmazonMQ, or SQS.
- Knowledge of microservice architecture, application design, distributed monitoring, SLOs, metrics, tracing, and log pipelines.
- Working knowledge of AWS, Linux, storage, networking, and PCI-scoped or equivalent environments.
- Strong technical writing, documentation, and communication skills; able to build consensus across teams in multiple time zones.
- Experience with highly trafficked web services.
Benefits
- Equity and 401(k)
- Medical, dental, and vision plans
- Company-paid short- and long-term disability coverage
- Paid time off, including flexible time off for exempt employees, vacation for non-exempt employees, and paid sick leave
- Paid parental leave
- Discounted meals and company perks