Esta vacante solo está disponible en inglés por ahora.

← Volver a Empleos
N

Senior Site Reliability Engineer

Navan·Israel·en
PresencialTiempo completoDevOps & SRETravel TechnologyEnterprise SoftwareFinTech

A travel technology company is seeking a hands-on Senior Site Reliability Engineer in Tel Aviv to operate reliable, scalable production platforms and support AI-powered applications. The role focuses on infrastructure, automation, observability, incident response, and dependable integration with AI

Responsabilidades

  • Support development and production operations for AI-powered travel applications.
  • Partner with teams integrating AI solutions, providers, and APIs, addressing reliability, authentication, quotas, rate limits, latency, and provider constraints.
  • Diagnose failures across AI workflows, provider APIs, configurations, permissions, responses, and related systems.
  • Implement and operate cloud infrastructure and reliable production platforms.
  • Build dashboards, alerts, traces, logs, and runbooks connected to SLOs and customer impact.
  • Prototype and productionize AI-assisted systems that improve operational effectiveness.
  • Create automation, tools, and workflows that reduce repetitive operational work and make operational knowledge easier to use.
  • Collaborate with product, platform, data, security, support, and incident response teams.

Requisitos

  • At least 5 years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
  • At least 3 years operating production, 24/7 customer-facing systems.
  • Hands-on experience delivering production infrastructure, platform tooling, and automation for engineering teams.
  • Strong software engineering skills in Python, Go, Java, or a similar language, including production-quality code, testing, monitoring, and documentation.
  • Experience with cloud infrastructure, container orchestration, Linux, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
  • Experience building, tuning, and automating observability systems using tools such as Grafana, Prometheus, New Relic, Datadog, or Splunk.
  • Familiarity with SLOs, incident response, on-call operations, root cause analysis, and blameless postmortems.
  • Ability to troubleshoot AI tools and provider or API issues, including rate limits, quotas, authentication, permissions, latency, contract changes, content quality, and service degradation.
  • Strong communication skills and the ability to collaborate with stakeholders and domain experts across the company.

Se valora

  • Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.

Compatibilidad

Más oportunidades

Vacantes similares

Nuevas vacantes en DevOps & SRE.

Senior Site Reliability Engineer – AI Platforms and Operations | CVZilla