Senior Site Reliability Engineer
Own the reliability and operational performance of critical production systems in a fully remote, globally distributed infrastructure team. Lead reliability targets, observability, incident response, capacity planning, resilient cloud infrastructure, safe deployments, and production tooling for a高?
Обязанности
- Own the reliability, instrumentation, operational performance, and real-world behavior of critical production systems.
- Define and maintain SLIs, SLOs, and error budgets, connecting reliability data to engineering priorities.
- Improve monitoring and alerting by increasing actionable signal and reducing noise.
- Lead high-severity incident response, service restoration, investigations, and actionable postmortems.
- Design and operate secure, resilient, and cost-efficient infrastructure with controlled failure modes and reduced blast radius.
- Perform load testing, profiling, saturation analysis, capacity planning, and proactive headroom management.
- Improve deployment safety through progressive delivery, automated rollback, and pre-production validation.
- Manage infrastructure and configuration through Terraform and related automation.
- Build production software and developer tooling that reduces operational toil.
- Run controlled failure testing, game days, and chaos exercises.
- Partner with engineering teams on production readiness, failure modes, rollback strategies, runbooks, and on-call handoffs.
- Participate in on-call rotations, improve escalation practices, and mentor engineers through reviews, pairing, and design feedback.
Требования
- 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering, preferably in cloud-based production environments.
- Experience owning systems end to end from design and implementation through production operation and improvement.
- Hands-on experience with SLIs, SLOs, error budgets, incident response, and postmortems.
- Strong knowledge of distributed-system failure modes, high-throughput and low-latency systems, capacity constraints, degradation, and load shedding.
- Strong cloud infrastructure fundamentals, including networking, load balancing, containerization, Kubernetes or EKS, and distributed systems.
- Extensive experience with Terraform or equivalent infrastructure-as-code technologies.
- Strong programming skills in Go, Python, or a comparable language, with experience shipping production software.
- Experience with observability platforms such as Datadog, Prometheus, Grafana, or OpenTelemetry, including direct instrumentation.
- Production experience with Redis or ElastiCache, including clustering, sharding, failover, eviction policies, and scaling.
- Strong software engineering practices covering source control, code review, testing, and safe deployments.
- Excellent written and verbal English communication skills for technical documents, reviews, incident communications, and postmortems.
- Comfort using AI tools in engineering workflows with sound judgment about their limitations.
Условия и преимущества
- Competitive compensation aligned with the hiring market and relevant experience.
- Fully remote working environment.
- Work with a globally distributed engineering organization.
- Significant ownership of production reliability, infrastructure, and operational practices.
- Work on high-throughput, low-latency systems with direct customer impact.
- Exposure to modern cloud infrastructure, observability, distributed systems, infrastructure as code, and AI-assisted engineering workflows.
- Opportunity to influence engineering standards and production-readiness practices.
- Collaborative, inclusive environment focused on ownership, autonomy, continuous improvement, and knowledge sharing.
- Candidates must be authorized to work from their home location; visa sponsorship is not provided.