Senior AI Infrastructure Engineer
Join a highly technical R&D organization to build and operate the infrastructure, tooling, and platforms that enable AI development in a secure, fully on-premises environment.
Обязанности
- Design, build, and maintain AI infrastructure and platforms.
- Develop and operate training and inference environments, including GPU clusters, schedulers, and containers.
- Build internal tools and services for model versioning, deployment, and monitoring.
- Collaborate with AI researchers and developers to optimize workflows and system performance.
- Work with DevOps teams on CI/CD pipelines, automation, and infrastructure as code.
- Ensure the reliability, reproducibility, and observability of AI systems.
- Troubleshoot complex issues across infrastructure, code, and runtime environments.
- Contribute to architecture decisions in a secure, fully on-premises environment.
Требования
- At least 4 years of experience in infrastructure engineering, DevOps, or backend systems.
- Strong experience with Linux systems and networking.
- Hands-on experience with Docker and Kubernetes.
- Experience building and maintaining CI/CD pipelines.
- Strong programming skills in Python, Go, or a similar language.
- Experience working with software engineers in production environments.
- Solid understanding of distributed systems and system design.
- Ability to work in a highly technical, fast-paced R&D environment.