Senior DevOps Engineer
Lead infrastructure and operations for a video architecture team, including on-prem and cloud compute, distributed pipelines for regression testing and ML workloads, and CI/CD environments. Hybrid role with 4 days in office.
Responsibilities
- Transform one-off research workflows into dependable, automated systems in collaboration with architects and algorithms engineers
- Stand up and operate on-prem GPU clusters, cloud bursts, queues, and schedulers such as Slurm and Kubernetes
- Develop decentralized workflows for extensive regression testing and experiments across hardware simulations and machine learning workloads
- Lead the team’s CI/CD and development environments, including container images and tooling
Requirements
- B.Sc. in Computer Science or Electrical/Computer Engineering
- 5+ years in a DevOps, SRE, MLOps, Research-Ops, or platform-engineering role
- Strong Linux fundamentals: shell, processes, networking, filesystems, systemd, performance tools
- Proficiency in Python for production-quality code
- Hands-on experience with a major cloud ecosystem (OCI, AWS, Azure, or GCP) and Infrastructure as Code (Terraform, Pulumi, or similar)
- Experience with containers and orchestration using Docker and Kubernetes or HPC schedulers like Slurm
- Experience designing and scaling CI/CD flows (GitLab CI, GitHub Actions, Jenkins) and operating distributed batch pipelines
Nice to have
- Familiarity with video compression or codecs (NVENC, NVDEC, FFmpeg, GStreamer)
- GPU-aware infrastructure experience: CUDA toolkit installs, driver versioning, MIG, NCCL
- Reading-level comfort with C++ to debug builds or trace benchmark issues
- Observability experience with Prometheus, Grafana, OpenTelemetry, and structured logging