Senior DevOps Engineer
Own infrastructure and deployment for real-time video and AI systems running primarily on customer premises and also in the cloud. Shape reliability, security, and engineering practices as an early member of the team.
Responsabilidades
- Design, build, and own production infrastructure for real-time, mission-critical systems.
- Design on-premises customer deployments, including installation, upgrades, remote management, and recovery.
- Build and maintain Dockerized services and multi-container deployments with Docker Compose.
- Manage cloud infrastructure for hosted environments.
- Own CI/CD pipelines, release processes, and deployment strategy from code commit to customer site.
- Build observability across metrics, logging, alerting, and on-call practices.
- Operate GPU workloads and ensure video and AI inference pipelines run reliably within hardware constraints.
- Own infrastructure-as-code, secrets management, networking, and security hardening.
- Contribute to backend services and APIs; improve data processing, database access, and real-time event handling.
- Partner with backend, ML, and product teams to make systems deployable, operable, and scalable.
- Set engineering standards for reliability, deployment, observability, security, and operational excellence.
- Contribute to the technical roadmap for systems operating in demanding real-world environments.
Requisitos
- At least 5 years in DevOps, SRE, platform, or infrastructure engineering, including backend development.
- Production experience with Docker and Docker Compose.
- Experience deploying and operating systems on physical servers at customer sites.
- Experience with at least one major cloud provider: AWS, GCP, or Azure.
- Experience with CI/CD pipelines and infrastructure-as-code tools such as Terraform or Ansible.
- Strong understanding of Linux, networking, and system troubleshooting.
- Experience with observability tools such as Prometheus, Grafana, or Loki.
- Ability to write production-quality backend code; Python preferred.
- Production PostgreSQL experience, including backups, replication, upgrades, and performance tuning.
- Strong understanding of system design, reliability, scalability, and security.
- Computer science degree or equivalent.
Se valora
- Built an on-premises environment from scratch, including hardware, networking, deployment, and operations.
- Experience with air-gapped or offline deployments and remote customer-site management.
- Experience with edge devices or fleet management, such as NVIDIA Jetson.
- Experience with GPU workloads, inference pipelines, or NVIDIA DeepStream.
- Experience with real-time video, streaming pipelines, or distributed camera networks.
- Strong Python experience with FastAPI or a similar framework.
- Experience with ML infrastructure, LLMs, vector databases, or AI agents.
- Experience designing systems under strict latency, bandwidth, or hardware constraints.
- Background in defense technology, security, robotics, autonomous systems, or physical-world AI.
- Early-stage startup experience.
Beneficios
- Work on real-time vision, multimodal AI, and distributed video systems.
- Help build technology intended to protect people and critical infrastructure.
- Take ownership of architecture, infrastructure, and technical direction from day one.
- Join an early-stage team building systems for demanding real-world environments.