MLOps Engineer
Own the production lifecycle of classical ML, LLM, and agentic AI solutions by building reusable deployment platforms and operating reliable model pipelines.
Responsibilities
- Design and build a reusable deployment platform for ML and AI
- Own model and pipeline packaging, deployment, monitoring, retraining, and incident response
- Implement CI/CD, monitoring, and observability for production models
- Evaluate AI solutions based on quality, latency, and cost
- Support reliable operation of classical ML, LLM-based, and agentic solutions
Requirements
- Hands-on experience deploying and operating AI solutions in production, including agentic pipelines
- Experience deploying and maintaining classical ML and LLM-based models in batch and online serving modes
- Experience with experiment tracking, model versioning, model registries, and deployment pipelines
- Strong Python and software engineering fundamentals
- Production experience with GCP, AWS, or Azure cloud platforms
- Ability to design generic, reusable deployment platforms rather than one-off deployments
Nice to have
- A degree in computer science, engineering, statistics, or a related field
- Production experience with Databricks and MLflow
- Experience with infrastructure as code, feature stores, streaming data, or real-time serving at scale
- Familiarity with agent orchestration and AI evaluation frameworks