Senior Site Reliability Engineer
A global technology organization is seeking an experienced Site Reliability Engineer to maintain and improve the reliability, stability, and performance of large-scale Linux production infrastructure.
Обязанности
- Administer, maintain, and troubleshoot large-scale Linux production environments
- Maintain high availability, reliability, and performance of critical infrastructure services
- Develop and maintain operational automation with Bash and Python
- Manage and optimize configuration management platforms
- Support Kubernetes and container-based infrastructure
- Troubleshoot issues across operating systems, networking, virtualization, and applications
- Manage DNS, DHCP, LDAP, and Active Directory integrations
- Maintain VMware and KVM virtualized environments
- Collaborate with engineering and infrastructure teams to improve operational efficiency
- Participate in production incident response and root cause analysis
Требования
- At least 4 years of hands-on experience managing Linux production environments
- Strong experience with RHEL, Rocky Linux, CentOS, Ubuntu, or similar distributions
- Expertise in Linux system administration and troubleshooting
- Experience with Bash and Python scripting
- Experience with Ansible, Puppet, Chef, or similar configuration management tools
- Experience supporting Kubernetes and containerized environments
- Strong understanding of TCP/IP, routing, and switching
- Experience with VMware, KVM, or similar virtualization technologies
- Ability to resolve complex infrastructure issues in enterprise-scale environments
Будет плюсом
- Experience in large-scale data center environments
- Experience supporting high-availability infrastructure
- Experience managing large server fleets
- Familiarity with monitoring and operational best practices