← Back to jobs
N

Senior Site Reliability Engineer

·Israel
Not specifiedFull-timeDevOps & SRESemiconductors

Senior Site Reliability Engineer responsible for designing and operating large-scale compute cloud infrastructure for chip design and deep learning workloads. Focus on architecture, resource optimization, capacity planning, and scaling across global HPC environments.

Responsibilities

  • Provide leadership in design and implementation of large-scale compute cloud for chip modelers, designers, and deep learning experts.
  • Identify architectural changes and innovative approaches in cloud architecture and design.
  • Address strategic challenges: resource utilization in heterogeneous compute environments, evolving private/public cloud strategy, capacity modeling, and multi-year scaling planning.

Requirements

  • B.Sc in Computer Science, Electrical Engineering or related field or equivalent experience.
  • 8+ years of experience designing and operating large scale compute infrastructure.
  • Experience with job schedulers (IBM/Platform LSF, SGE, SLURM, Marathon, Chronos).
  • Solid understanding of cluster configuration management tools (Ansible, Puppet, Chef, Salt).
  • Experience providing compute services using public cloud (AWS, Azure, Google Cloud).
  • Strong script-writing skills: Python, Bash, Perl.
  • Knowledge of deploying PaaS microservices: Docker, Docker Swarm, Kubernetes.
  • Understanding of fast distributed and network attached storage solutions and Linux file systems; ability to recommend and implement OS performance/reliability improvements.

Nice to have

  • Linux certification from well-known vendor (RedHat, Oracle, etc.).
  • Prior experience managing large-scale Kubernetes deployment in production.
  • Strong skills in modern container networking and storage architecture.
  • Well-known Cloud Certification(s).

Relevance

More opportunities

Similar jobs

Finding the best alternatives for you…

Questions, answered

Frequently asked questions