Senior Site Reliability Engineer – Storage
Own the reliability, performance, security, and scalability of global NAS, SAN, and object storage platforms. Design storage solutions, automate infrastructure operations, improve observability, and lead incident response for highly available services.
תחומי אחריות
- Design, deploy, and operate production NAS, SAN, and object storage platforms
- Capture requirements, architect storage solutions, and drive end-to-end implementation
- Develop automation for provisioning, configuration, monitoring, incident response, and lifecycle management
- Participate in on-call rotations, troubleshoot complex storage issues, and lead root cause analysis
- Define and track SLOs, SLIs, and error budgets using observability and analytics
- Analyze capacity and usage trends, forecast demand, and recommend scaling or optimization strategies
- Maintain runbooks, standard operating procedures, and service documentation
- Collaborate with SRE, infrastructure, networking, and application teams across distributed operations
- Mentor junior engineers and promote SRE practices
דרישות
- 12+ years of experience in site reliability, DevOps, or infrastructure engineering, with substantial storage expertise
- Bachelor’s degree in computer science, computer engineering, or a related technical field, or equivalent practical experience
- Hands-on experience designing, deploying, and operating enterprise NAS, SAN, and/or object storage platforms
- Strong knowledge of SRE practices, including SLOs, SLIs, error budgets, incident management, observability, and postmortems
- Proficiency with infrastructure as code and configuration management tools such as Terraform, Ansible, Puppet, or SaltStack
- Experience operating highly available and scalable infrastructure with provisioning, monitoring, and remediation automation
- Experience with containers, virtualization, CI/CD, and version control systems
- Strong scripting or programming skills in Python, Go, Shell, or similar languages
יתרון
- Storage experience supporting high-performance computing, AI/ML workloads, or large-scale analytics
- Experience debugging distributed systems and storage performance issues
- Track record of data-driven reliability improvements through automation
- Experience leading technical initiatives or mentoring engineers
הטבות
- Competitive salary
- Generous benefits package