Senior Site Reliability Engineer
A travel technology company is seeking a hands-on Senior Site Reliability Engineer in Tel Aviv to operate reliable, scalable production platforms and support AI-powered applications. The role focuses on infrastructure, automation, observability, incident response, and dependable integration with AI
Responsabilidades
- Support development and production operations for AI-powered travel applications.
- Partner with teams integrating AI solutions, providers, and APIs, addressing reliability, authentication, quotas, rate limits, latency, and provider constraints.
- Diagnose failures across AI workflows, provider APIs, configurations, permissions, responses, and related systems.
- Implement and operate cloud infrastructure and reliable production platforms.
- Build dashboards, alerts, traces, logs, and runbooks connected to SLOs and customer impact.
- Prototype and productionize AI-assisted systems that improve operational effectiveness.
- Create automation, tools, and workflows that reduce repetitive operational work and make operational knowledge easier to use.
- Collaborate with product, platform, data, security, support, and incident response teams.
Requisitos
- At least 5 years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
- At least 3 years operating production, 24/7 customer-facing systems.
- Hands-on experience delivering production infrastructure, platform tooling, and automation for engineering teams.
- Strong software engineering skills in Python, Go, Java, or a similar language, including production-quality code, testing, monitoring, and documentation.
- Experience with cloud infrastructure, container orchestration, Linux, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
- Experience building, tuning, and automating observability systems using tools such as Grafana, Prometheus, New Relic, Datadog, or Splunk.
- Familiarity with SLOs, incident response, on-call operations, root cause analysis, and blameless postmortems.
- Ability to troubleshoot AI tools and provider or API issues, including rate limits, quotas, authentication, permissions, latency, contract changes, content quality, and service degradation.
- Strong communication skills and the ability to collaborate with stakeholders and domain experts across the company.
Se valora
- Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.