Site Reliability Engineer
Own production health and performance for an industrial AI data quality platform. Guide customers through installation, manage high-throughput data pipelines and ML systems. Provide second-line support and contribute to DevOps with Python, Kubernetes, AWS, and Terraform.
Responsibilities
- Own production end-to-end: installations, maintenance, migrations, monitoring
- Manage high-throughput data pipelines and cutting-edge ML systems
- Guide customers through technical requirements and integrations
- Collaborate with Sales and Customer Success for technical clarity
- Act as second-line technical support for complex issues
- Contribute to DevOps: CI pipelines, infrastructure, deployments
Requirements
- Works autonomously and learns independently
- Good written and spoken English
- Clear communication with customers and stakeholders
- Proficiency in Python and self-managed services, databases, custom tooling
- Familiarity with troubleshooting networked environments (DNS, firewalls, web proxies)