Senior Site Reliability Engineer
Software engineer with 2+ years of experience in large-scale distributed systems and infrastructure. Proficient in Go, Python, and cloud technologies. Focus on automating operational workflows, enhancing system reliability, and implementing robust monitoring and alerting solutions. Skilled in Kubernetes, Terraform, and incident management.
Responsibilities
- Design, build, and maintain software and systems to enhance reliability, availability, and performance.
- Create and improve software tools and platforms that automate operational workflows and reduce manual toil at scale.
- Develop and manage infrastructure configurations using Go for secure, consistent, and auditable management of cloud resources.
- Design and implement sophisticated monitoring, logging, and tracing solutions.
- Analyze past incidents and develop software solutions to prevent recurrence; build automation for incident detection, diagnosis, and resolution.
Requirements
- Bachelor’s degree or equivalent practical experience.
- 2+ years of software development experience, or 1 year with an advanced degree.
- 2+ years of experience with large-scale infrastructure, distributed systems, or networks, or experience with compute technologies, storage, or hardware architecture.
- Experience in one or more programming languages such as Go, Python, Java, C++, or C#.
Nice to have
- Experience in designing, analyzing, and troubleshooting large-scale distributed systems.
- Experience with container orchestration (e.g., Kubernetes, GKE).
- Experience with infrastructure automation and configuration management tools (e.g., Terraform, Ansible).
- Experience with incident management and on-call rotations.
- Experience with security principles and practices, including access management and compliance.
- Experience with monitoring, logging, and alerting best practices and tools.