Principal Site Reliability Engineer
Lead production infrastructure, platforms, and services, driving scalable solutions, system architecture, and reliability improvements. Collaborate across teams to enhance troubleshooting, automation, and incident management.
Responsibilities
- Lead production infrastructure and platform maintenance, designing highly scalable environments.
- Collaborate with team leads to direct work and adapt solutions for organizational needs.
- Proactively design new scalable solutions for configuration deployments, monitoring, and logging.
- Lead deep drill-down analysis into production support operations and provide necessary fixes.
- Investigate errors, debug source code and performance bottlenecks, and determine root cause analysis needs.
- Build complex enhancements such as automating alerts and review proposed enhancements to resolve defects.
- Drive design, development, configuration, and deployment of enterprise supporting systems.
- Lead code and architecture reviews, resolving complex performance and scalability issues.
- Identify opportunities to improve health, scalability, and resiliency of products and services.
- Manage and coordinate complex tasks, monitoring timelines and deliverables across projects.
- Coach and mentor junior team members, fostering continuous learning and knowledge sharing.
Requirements
- 10+ years of experience in server architecture, system administration, software development, or cloud application delivery.
- Proficiency in writing maintainable, effective code in high-level programming languages.
- Experience with configuration management and deployment tools (e.g., Kubernetes, Terraform, Ansible, Chef, Puppet).
- Ability to architect end-to-end solutions meeting business and technical requirements.
- Knowledge of cloud architecture, including designing scalable, reliable, and performant cloud services.
- Demonstrated ability to conduct root cause analysis and implement corrective action plans.
- Experience with Linux system administration, networking, storage, compute, and virtualization.
- Experience participating in or running incident bridges of significant scale.
- Experience in SRE, cloud technical support, cloud operations, and incident management.
- Customer focus with a passion for delighting customers.
- Demonstrated ability to quickly learn new technical disciplines and train/mentor others.
- Israeli nationals only; willingness to obtain and maintain security clearance for public sector projects.
Benefits
- Competitive benefits package
- Flexible medical insurance
- Life insurance
- Retirement options
- Volunteer programs