Production Engineer
Ensure the reliability, performance, and operational excellence of a distributed federated computing platform supporting AI/ML research in regulated industries. Focus on production operations, monitoring, incident response, and automation. 3-5 years of DevOps/SRE experience required.
Responsibilities
- Maintain and support production environments across cloud and on-premises deployments.
- Develop and improve monitoring, alerting, and logging systems.
- Investigate, troubleshoot, and resolve complex infrastructure and application issues.
- Manage production deployments, upgrades, and maintenance activities.
- Identify operational bottlenecks and improve reliability, scalability, and automation.
- Collaborate with Backend, DevOps, and Product Engineering teams.
- Contribute to internal tooling and automation for deployment and support efficiency.
Requirements
- 3-5 years of experience in Production Engineering, DevOps, SRE, or similar roles.
- Experience with cloud environments (AWS and/or GCP).
- Experience operating Kubernetes-based environments.
- Strong Linux administration and troubleshooting skills.
- Experience with Docker and containerized workloads.
- Proficiency in Python and scripting for automation.
- Experience with Infrastructure-as-Code (Terraform, Ansible).
- Experience with monitoring and observability systems (Prometheus, Grafana).
- Experience with CI/CD tools such as GitHub Actions.
- Experience troubleshooting networking issues in distributed environments.
- Familiarity with GitOps concepts and tools (ArgoCD preferred).
Nice to have
- Experience supporting AI/ML products or platforms.
- Experience operating distributed systems.
- Experience supporting customer-facing production systems.
- Experience with security-focused or privacy-sensitive workloads.
- Experience with VPN technologies, mTLS, ingress systems, or service networking.