Senior DevOps Engineer
Own the infrastructure, delivery, security, and reliability lifecycle for an AI observability platform deployed in SaaS and customer-managed environments.
Responsibilities
- Own DevOps, infrastructure, security, release, and reliability from development environments and CI/CD through deployment and ongoing operations
- Build and operate Kubernetes environments and Helm deployments for a Python, FastAPI, Node.js, and React microservices platform across SaaS and on-premises environments
- Develop GitLab CI/CD pipelines with automated testing, container builds, vulnerability scanning, and versioned image and Helm chart publication through JFrog Artifactory
- Automate infrastructure provisioning, configuration, upgrades, and routine operational workflows
- Operate PostgreSQL, Temporal services, and S3-compatible storage, including capacity planning, backups, recovery testing, and safe migrations
- Improve release reliability through deployment validation, reduced-downtime strategies, persistent-state protection, and recovery plans for active workflows
- Implement observability and security practices using OpenTelemetry, Datadog/Grafana, Langfuse, secrets management, identity integration, TLS, Kubernetes RBAC, network policies, and container hardening
- Partner with software and AI engineers to troubleshoot distributed systems, investigate incidents, define reliability targets, and improve platform performance and resource efficiency
Requirements
- At least 5 years in DevOps, site reliability, or platform engineering for distributed applications and microservices
- Hands-on Kubernetes, Docker, and Helm experience, including networking, storage, scheduling, scaling, and troubleshooting
- Linux administration skills and Python and Bash automation experience
- Experience with infrastructure as code and configuration tools such as Terraform and Ansible
- Experience building and maintaining CI/CD pipelines, runners, container registries, artifact management, quality gates, and secure release practices
- Experience operating PostgreSQL or a comparable relational database, including SQL, migrations, backup and restore, and performance troubleshooting
- Strong networking and observability fundamentals, including TCP/IP, DNS, HTTP, TLS, load balancing, ingress, metrics, logs, traces, dashboards, and alerting
- Understanding of secure infrastructure operations and incident response; bachelor's degree in computer science, software engineering, or a related field, or equivalent experience
Nice to have
- Experience operating AI applications, agent platforms, or LLM services
- Familiarity with Temporal, LangGraph, MCP, Langfuse, ClickHouse, Redis/Valkey, or S3-compatible storage
- Experience with OpenTelemetry, Datadog APM, or Prometheus/Grafana
- Experience with self-hosted Kubernetes, OpenShift, Kubernetes operators, CloudNativePG, GPU clusters, or AI data centers
- Experience building AMD64 and ARM64 container images, optimizing BuildKit pipelines, or securing software supply chains
Benefits
- Competitive salary
- Generous benefits package