Site Reliability Engineer
We are looking for a Site Reliability Engineer to lead infrastructure projects, manage Kubernetes environments, and enhance reliability across cloud platforms. You will work with cutting-edge technologies like Kafka, Istio, and OpenTelemetry to drive automation and scalability.
Responsibilities
- Act as a hands-on technical leader for cloud infrastructure.
- Collaborate cross-functionally to design scalable solutions.
- Drive long-term infrastructure projects from design to implementation.
- Improve system reliability, performance, and cost-efficiency at scale.
- Manage and scale Kubernetes clusters and observability agents.
- Write and maintain Kubernetes Controllers.
Requirements
- 5+ years of experience in DevOps, SRE, platform engineering, or infrastructure roles.
- Deep understanding of Kubernetes: API, CNI, scheduling, container runtimes.
- Strong hands-on experience with Kafka and Istio (or similar technologies), and core networking protocols (HTTP, gRPC, TLS).
- Proven experience managing large-scale cloud infrastructure (AWS, GCP, etc.).
- Experience in incident response and troubleshooting complex distributed systems.
- Some software engineering experience, preferably in Golang.
- Passion for automation, performance tuning, and operational excellence.