← Back to jobs
O

Principal Site Reliability Engineer

·Israel
On-siteFull-timeDevOps & SREEnterprise SoftwareCloud Computing

Lead production infrastructure, platforms, and services, driving scalable solutions, system architecture, and reliability improvements. Collaborate across teams to enhance troubleshooting, automation, and incident management.

Responsibilities

  • Lead production infrastructure and platform maintenance, designing highly scalable environments.
  • Collaborate with team leads to direct work and adapt solutions for organizational needs.
  • Proactively design new scalable solutions for configuration deployments, monitoring, and logging.
  • Lead deep drill-down analysis into production support operations and provide necessary fixes.
  • Investigate errors, debug source code and performance bottlenecks, and determine root cause analysis needs.
  • Build complex enhancements such as automating alerts and review proposed enhancements to resolve defects.
  • Drive design, development, configuration, and deployment of enterprise supporting systems.
  • Lead code and architecture reviews, resolving complex performance and scalability issues.
  • Identify opportunities to improve health, scalability, and resiliency of products and services.
  • Manage and coordinate complex tasks, monitoring timelines and deliverables across projects.
  • Coach and mentor junior team members, fostering continuous learning and knowledge sharing.

Requirements

  • 10+ years of experience in server architecture, system administration, software development, or cloud application delivery.
  • Proficiency in writing maintainable, effective code in high-level programming languages.
  • Experience with configuration management and deployment tools (e.g., Kubernetes, Terraform, Ansible, Chef, Puppet).
  • Ability to architect end-to-end solutions meeting business and technical requirements.
  • Knowledge of cloud architecture, including designing scalable, reliable, and performant cloud services.
  • Demonstrated ability to conduct root cause analysis and implement corrective action plans.
  • Experience with Linux system administration, networking, storage, compute, and virtualization.
  • Experience participating in or running incident bridges of significant scale.
  • Experience in SRE, cloud technical support, cloud operations, and incident management.
  • Customer focus with a passion for delighting customers.
  • Demonstrated ability to quickly learn new technical disciplines and train/mentor others.
  • Israeli nationals only; willingness to obtain and maintain security clearance for public sector projects.

Benefits

  • Competitive benefits package
  • Flexible medical insurance
  • Life insurance
  • Retirement options
  • Volunteer programs

Relevance

More opportunities

Similar jobs

Finding the best alternatives for you…

Questions, answered

Frequently asked questions