Senior Site Reliability Engineer
Join an advertising technology company’s R&D Infrastructure team in Tel Aviv to build, scale, and operate highly available hybrid infrastructure across on-premise systems, public cloud, and AI/ML Kubernetes environments. The role is hybrid, with three days in the office.
Responsibilities
- Operate and improve highly available, performant, and cost-efficient hybrid infrastructure spanning on-premise systems, public cloud, and AI/ML clusters
- Build internal software tooling and manage Infrastructure as Code pipelines
- Automate repetitive operational work using Go, Python, or Rust
- Troubleshoot issues across CDN configuration, Linux kernel performance, and network layers
- Design and maintain monitoring, telemetry, and alerting systems
- Participate in on-call rotations, lead incident resolution, and conduct blameless post-mortems
Requirements
- At least 7 years of experience managing, scaling, and troubleshooting large-scale distributed Linux environments in production
- Deep knowledge of Linux system internals and network protocols including TCP/IP, DNS, HTTP, and gRPC
- Hands-on experience with edge and CDN services such as Fastly, Cloudflare, Akamai, or CloudFront
- Experience with Infrastructure as Code and orchestration tools such as Terraform, Ansible, Puppet, ArgoCD, or Jenkins
- Production experience managing containerized environments with Kubernetes and Docker
- Strong programming skills in Go, Python, or Rust
Nice to have
- Experience designing and operating telemetry, metrics, and alerting systems at scale
- Experience with Prometheus, Grafana, ELK, or comparable logging and observability tools
- Experience optimizing infrastructure costs and resource efficiency across cloud and on-premise environments
Benefits
- Comprehensive health and other employee benefits
- Fully stocked kitchen
- Gym partnerships, parking, and location-specific perks
- Hybrid work schedule with flexibility
- Inclusive and equal-opportunity workplace