← Back to jobs
I

Senior Platform Reliability Engineer

·Israel
RemoteFull-timeBackend DevelopmentSoftware Development

Join a reliability team that engineers incidents out of existence for an AI-native community and social engagement platform serving Fortune 100 brands. You will carry production, build autonomous AI-driven operational workflows, and drive systemic prevention in a fully remote, globally distributed environment.

Responsibilities

  • Own production incident response during your shift, commanding diagnosis, mitigation, and service restoration with personal accountability for customer impact.
  • Engineer autonomous operational workflows by building, deploying, and refining AI agents that handle pre-triage, change-gate validation, auto-healing, RCA drafting, and preventive-fix tracking.
  • Ship safe production changes through quality gates with validated rollback plans, aborting immediately when telemetry deviates.
  • Investigate incidents to true root cause, identify systemic preventions, build them, and track to production closure.
  • Generalize every manual intervention into agents, runbooks, or guardrails to expand the autonomous layer.
  • Multiply team knowledge by encoding procedures, context, and decision logic for agent retrieval and responder readiness.

Requirements

  • 5+ years of hands-on production operations in SRE, Platform Engineering, DevOps, or Cloud Infrastructure at SaaS scale with real first-responder incident experience.
  • Battle-tested on AWS in multi-AZ, multi-account environments with infrastructure-as-code, production incident management, and change control with gates and rollbacks.
  • Radically self-directed with the ability to identify gaps, prioritize, ship solutions, and challenge standards without waiting for direction.
  • AI-native in practice, routinely delegating operational work to agents, critically evaluating output, and iterating on capabilities using tools like Claude Code, Codex, or custom frameworks.
  • AWS Solutions Architect – Associate or higher, or equivalent production track record.
  • Fluent, precise English for incident-bridge communication and long-form writing.
  • Committed to shift-based coverage with required time-zone overlap.
  • Residence in an OFAC-clear country.

Nice to have

  • Original contributions to agentic operations or AIOps, such as open-source tooling, technical writing, conference talks, or shipped internal platforms.
  • Production experience with multi-tenant B2B SaaS, community platforms, social tools, customer-experience products, or observability systems.
  • Working knowledge of Grafana, Prometheus, Datadog, PagerDuty, or OpsGenie, plus Azure exposure alongside AWS depth.
  • Evidence of deep, sustained obsession with a hard problem, professional or personal.

Benefits

  • Build the future of AI-driven reliability on a platform depended on by Fortune 100 brands.
  • Develop agent-ops skills, incident patterns, and architectural instincts that are industry-defining.
  • Work with enterprise clients at startup velocity with weekly delivery cycles and fast decision-making.
  • Uncapped tooling and compute investment in the agent harness.
  • Fully remote, global team.

Relevance

More opportunities

Similar jobs

Finding the best alternatives for you…

Questions, answered

Frequently asked questions