Senior Platform Reliability Engineer
Join a reliability team that engineers incidents out of existence for an AI-native community and social engagement platform serving Fortune 100 brands. You will carry production, build autonomous AI-driven operational workflows, and drive systemic prevention in a fully remote, globally distributed environment.
Responsibilities
- Own production incident response during your shift, commanding diagnosis, mitigation, and service restoration with personal accountability for customer impact.
- Engineer autonomous operational workflows by building, deploying, and refining AI agents that handle pre-triage, change-gate validation, auto-healing, RCA drafting, and preventive-fix tracking.
- Ship safe production changes through quality gates with validated rollback plans, aborting immediately when telemetry deviates.
- Investigate incidents to true root cause, identify systemic preventions, build them, and track to production closure.
- Generalize every manual intervention into agents, runbooks, or guardrails to expand the autonomous layer.
- Multiply team knowledge by encoding procedures, context, and decision logic for agent retrieval and responder readiness.
Requirements
- 5+ years of hands-on production operations in SRE, Platform Engineering, DevOps, or Cloud Infrastructure at SaaS scale with real first-responder incident experience.
- Battle-tested on AWS in multi-AZ, multi-account environments with infrastructure-as-code, production incident management, and change control with gates and rollbacks.
- Radically self-directed with the ability to identify gaps, prioritize, ship solutions, and challenge standards without waiting for direction.
- AI-native in practice, routinely delegating operational work to agents, critically evaluating output, and iterating on capabilities using tools like Claude Code, Codex, or custom frameworks.
- AWS Solutions Architect – Associate or higher, or equivalent production track record.
- Fluent, precise English for incident-bridge communication and long-form writing.
- Committed to shift-based coverage with required time-zone overlap.
- Residence in an OFAC-clear country.
Nice to have
- Original contributions to agentic operations or AIOps, such as open-source tooling, technical writing, conference talks, or shipped internal platforms.
- Production experience with multi-tenant B2B SaaS, community platforms, social tools, customer-experience products, or observability systems.
- Working knowledge of Grafana, Prometheus, Datadog, PagerDuty, or OpsGenie, plus Azure exposure alongside AWS depth.
- Evidence of deep, sustained obsession with a hard problem, professional or personal.
Benefits
- Build the future of AI-driven reliability on a platform depended on by Fortune 100 brands.
- Develop agent-ops skills, incident patterns, and architectural instincts that are industry-defining.
- Work with enterprise clients at startup velocity with weekly delivery cycles and fast decision-making.
- Uncapped tooling and compute investment in the agent harness.
- Fully remote, global team.