Senior Site Reliability Engineer
Responsible for operating production environments, automation, and optimization. Requires deep expertise in Linux, cloud infrastructure, and incident management, with ability to lead and mentor junior engineers.
Responsibilities
- Develop automation and optimization focused on operational excellence.
- Perform root cause analysis and resolve systemic issues.
- Install, monitor, maintain, support, and optimize production server hardware and software.
- Provide escalated technical support for complex issues.
- Coordinate support cases and lead internal resources to resolution.
- Assist with server OS and application upgrades, bug fixes, and patching.
- Lead communications with partners.
- Provide technical guidance and leadership to junior members.
- Participate in 24/7 support rotation.
Requirements
- Israeli nationality required.
- Willingness to obtain and maintain security clearance.
- Experience with Linux System Administration, Networking, Storage, Compute, and Virtualization.
- Understanding of programming languages such as Python, Java, Rust, or Go.
- Experience with technologies like Kubernetes, Terraform, Ansible, Chef, and Puppet.
- Experience in running incident bridges.
- Customer focus.
- Experience in SRE, cloud technical support, cloud operations, NOC or similar.
- Ability to quickly learn new technical disciplines.
Benefits
- Flexible medical insurance
- Life insurance
- Retirement options
- Volunteer programs