Junior Site Reliability Engineer
Join a health and human services company as a Junior Site Reliability Engineer. You will monitor and maintain high system reliability, working with engineering, product, and data teams to ensure performance, availability, and observability of production systems.
Responsibilities
- Monitor system performance, uptime, and other KPIs to ensure high availability.
- Design, implement, and maintain scalable and reliable infrastructure.
- Develop automation tools to reduce manual effort and improve system efficiency.
- Drive incident response, root cause analysis, and postmortems.
- Build and maintain CI/CD pipelines and infrastructure-as-code.
- Collaborate with development teams to ensure best practices in service design and deployment.
- Implement and advocate for SLOs, SLIs, and SLAs across services.
- Continuously improve observability, including logging, tracing, and metrics collection.
Requirements
- Bachelor's degree in Computer Science, completion of a DevOps course, or proven experience.
- Proven knowledge and hands-on experience with Docker.
- Proficiency in Node.js (JavaScript) and React.
- Basic experience or exposure to AWS cloud services.
- Knowledge of Kubernetes and containerized environments.
- Solid understanding of Linux/Unix systems and networking fundamentals.
- Proficiency in at least one programming or scripting language (e.g., Python, Go, Bash).
- Experience working with SQL and relational databases.
- Familiarity with monitoring and observability tools (e.g., Prometheus, Grafana, ELK, Datadog).