Senior Site Reliability Engineer
Senior Site Reliability Engineer responsible for designing and operating large-scale compute cloud infrastructure for chip design and deep learning workloads. Focus on architecture, resource optimization, capacity planning, and scaling across global HPC environments.
Responsibilities
- Provide leadership in design and implementation of large-scale compute cloud for chip modelers, designers, and deep learning experts.
- Identify architectural changes and innovative approaches in cloud architecture and design.
- Address strategic challenges: resource utilization in heterogeneous compute environments, evolving private/public cloud strategy, capacity modeling, and multi-year scaling planning.
Requirements
- B.Sc in Computer Science, Electrical Engineering or related field or equivalent experience.
- 8+ years of experience designing and operating large scale compute infrastructure.
- Experience with job schedulers (IBM/Platform LSF, SGE, SLURM, Marathon, Chronos).
- Solid understanding of cluster configuration management tools (Ansible, Puppet, Chef, Salt).
- Experience providing compute services using public cloud (AWS, Azure, Google Cloud).
- Strong script-writing skills: Python, Bash, Perl.
- Knowledge of deploying PaaS microservices: Docker, Docker Swarm, Kubernetes.
- Understanding of fast distributed and network attached storage solutions and Linux file systems; ability to recommend and implement OS performance/reliability improvements.
Nice to have
- Linux certification from well-known vendor (RedHat, Oracle, etc.).
- Prior experience managing large-scale Kubernetes deployment in production.
- Strong skills in modern container networking and storage architecture.
- Well-known Cloud Certification(s).