On-Premise Site Reliability Engineer
Own the reliability, automation, and deployment of an AI and cyber-defense platform across customer sites. Work hands-on with Kubernetes, bare-metal infrastructure, Linux, and hybrid cloud while partnering with Product, R&D, Architecture, and customer technical teams.
תחומי אחריות
- Lead end-to-end on-premise and hybrid deployments in collaboration with internal and customer technical teams.
- Deploy, configure, operate, and improve enterprise Kubernetes environments.
- Design, modify, and manage Helm charts across deployment environments.
- Build and maintain Ansible automation that makes operational tasks repeatable and reliable.
- Use Git to manage infrastructure configurations, manifests, and automation playbooks.
- Manage AWS components supporting hybrid deployments, staging, and cloud-to-on-premise data flows.
- Provide tier-three technical escalation for deployment, Linux networking, and Kubernetes issues.
- Improve delivery pipelines and bootstrap processes.
- Create and maintain technical documentation.
- Act as the technical authority for customer deployments and platform reliability.
דרישות
- 3–5 years of experience in enterprise infrastructure deployment, systems engineering, or on-premise operational reliability.
- Production-grade Kubernetes architecture, deployment, troubleshooting, and CNI networking experience.
- Experience creating, maintaining, and deploying Helm charts.
- Experience writing scalable Ansible roles and playbooks for configuration management, automation, and infrastructure provisioning.
- Working knowledge of Git and AWS, including EC2 and S3.
- Strong Ubuntu/Linux systems knowledge, container runtimes, and distributed application troubleshooting.
- Hands-on knowledge of routing, firewalls, and switching topologies, mainly Cisco.
- Working knowledge of iSCSI, SAN, local NVMe, enterprise storage arrays, and Kubernetes persistent volumes.
- Working knowledge of GPU-enabled Kubernetes nodes, NVIDIA drivers and runtimes, and troubleshooting.
- Experience deploying and maintaining software in air-gapped or offline environments, including registry mirroring and artifact staging.
- High-level English proficiency.
- Willingness to travel to customer sites for physical staging and deployments at least 30% of the time.
יתרון
- EU or additional citizenship.
- Experience with MongoDB, PostgreSQL, Neo4j, or RabbitMQ.
- Experience with Jira or Monday.
- Valid Israeli security clearance.