Principal Staff Site Reliability Engineer
Lead the evolution of global core infrastructure services across on-premises and cloud environments, improving reliability, performance, automation, and operational efficiency.
Responsabilidades
- Lead architecture initiatives and develop new infrastructure service offerings
- Design, scale, and deploy DNS, NTP/PTP, DHCP, and LDAP services
- Own infrastructure performance, reliability, automation, monitoring, availability, capacity planning, and lifecycle management
- Define service-efficiency metrics and drive software and hardware optimization
- Apply eBPF and XDP for observability, performance analysis, and DDoS mitigation
- Analyze system and capacity data and coordinate enterprise-wide capacity planning
- Develop tools for data collection, analysis, visualization, reporting, alerting, and monitoring
- Collaborate with engineering, product, program, and company leadership on infrastructure services
Requisitos
- Bachelor’s degree in engineering, computer science, mathematics, or a related field, or equivalent experience
- 12+ years of compute platform engineering experience
- Experience with large-scale automation and technical leadership
- Experience designing containerization architectures and distributed systems infrastructure
- Strong analytical skills and experience defining performance metrics
- Experience developing data analysis and performance profiling tools
- Proficiency in Go, Python, or a similar language
- Linux and kernel internals proficiency
- Experience operating bare-metal build infrastructure at scale
- Understanding of VLAN, VXLAN, SDN, BGP, and Anycast
Se valora
- Deep knowledge of DNS, LDAP, and security tools
- Hands-on container implementation experience
- Experience managing DNS and LDAP services at scale
- Understanding of microservices, infrastructure as code, and configuration management
Beneficios
- Competitive salary
- Comprehensive benefits package
- Equal opportunity workplace