Senior Site Reliability Engineer
Help maintain reliable production services by coordinating development, operations, maintenance teams, and external vendors.
Responsibilities
- Act as the central point of contact between development teams, operations, and external vendors.
- Investigate and resolve complex production incidents.
- Build and maintain monitoring and alerting systems.
- Use data-driven root-cause analysis to guide reliability and performance improvements.
- Apply AI tools to accelerate troubleshooting, automation, and reliability workflows.
Requirements
- Experience integrating across development, infrastructure, operations, and vendors.
- Hands-on experience with Splunk and Dynatrace for production troubleshooting.
- Experience configuring monitoring systems, dashboards, and alerts.
- Database skills with MongoDB, DocumentDB, and SQL.
- Familiarity with OCP and AWS environments.
- Strong technical drive, self-learning ability, ownership, and organizational skills.
- Collaborative approach, service orientation, and commitment to operational excellence and continuous improvement.
Nice to have
- Familiarity with AI development and operations tools such as Claude or Kiro.