Mid-level DevOps Engineer, Enterprise GenAI
Own DevOps practices for an enterprise generative AI platform, enabling reliable development, deployment, monitoring, scaling, and adoption of production GenAI services.
Responsabilidades
- Design, build, and maintain CI/CD pipelines for GenAI services, agents, and platforms.
- Operate and scale containerized GenAI workloads across Azure and AWS environments.
- Ensure the availability, reliability, observability, and performance of AI services.
- Manage repeatable, version-controlled infrastructure provisioning with infrastructure as code.
- Collaborate with AI, software, infrastructure, and cybersecurity teams on governed GenAI deployments.
- Support GenAI solutions throughout experimentation, development, production deployment, and operation.
- Implement monitoring, logging, versioning, and rollback mechanisms for GenAI solutions, MCPs, agents, and pipelines.
- Support LLM serving infrastructure, including GPU workload scheduling and resource management.
- Promote DevOps practices, automation, documentation, and ownership across the organization.
- Create and maintain runbooks, architecture diagrams, and operational playbooks.
Requisitos
- At least 3 years of hands-on DevOps engineering experience.
- Strong Linux and scripting experience, preferably with Python.
- Hands-on experience with Docker and Kubernetes.
- Experience designing and operating production CI/CD pipelines.
- Hands-on experience with infrastructure as code using Terraform, Pulumi, or equivalent.
- Strong problem-solving skills, ownership, and ability to work independently.
Se valora
- Experience with vector databases such as Elasticsearch, Weaviate, or pgvector in RAG or GenAI pipelines.
- Experience supporting AI, ML, or GenAI systems in production.
- Experience with Azure and hybrid or on-premises environments.
- Familiarity with monitoring, logging, and observability tools.
- Experience in regulated, security-sensitive enterprise environments.