← Back to jobs
N

Senior DevOps Engineer

NVIDIA·Ra'anana, Israel·en
Not specifiedFull-timeDevOps & SRESemiconductorsComputer Hardware

Own the infrastructure, delivery, security, and reliability lifecycle for an AI observability platform deployed in SaaS and customer-managed environments.

Responsibilities

  • Own DevOps, infrastructure, security, release, and reliability from development environments and CI/CD through deployment and ongoing operations
  • Build and operate Kubernetes environments and Helm deployments for a Python, FastAPI, Node.js, and React microservices platform across SaaS and on-premises environments
  • Develop GitLab CI/CD pipelines with automated testing, container builds, vulnerability scanning, and versioned image and Helm chart publication through JFrog Artifactory
  • Automate infrastructure provisioning, configuration, upgrades, and routine operational workflows
  • Operate PostgreSQL, Temporal services, and S3-compatible storage, including capacity planning, backups, recovery testing, and safe migrations
  • Improve release reliability through deployment validation, reduced-downtime strategies, persistent-state protection, and recovery plans for active workflows
  • Implement observability and security practices using OpenTelemetry, Datadog/Grafana, Langfuse, secrets management, identity integration, TLS, Kubernetes RBAC, network policies, and container hardening
  • Partner with software and AI engineers to troubleshoot distributed systems, investigate incidents, define reliability targets, and improve platform performance and resource efficiency

Requirements

  • At least 5 years in DevOps, site reliability, or platform engineering for distributed applications and microservices
  • Hands-on Kubernetes, Docker, and Helm experience, including networking, storage, scheduling, scaling, and troubleshooting
  • Linux administration skills and Python and Bash automation experience
  • Experience with infrastructure as code and configuration tools such as Terraform and Ansible
  • Experience building and maintaining CI/CD pipelines, runners, container registries, artifact management, quality gates, and secure release practices
  • Experience operating PostgreSQL or a comparable relational database, including SQL, migrations, backup and restore, and performance troubleshooting
  • Strong networking and observability fundamentals, including TCP/IP, DNS, HTTP, TLS, load balancing, ingress, metrics, logs, traces, dashboards, and alerting
  • Understanding of secure infrastructure operations and incident response; bachelor's degree in computer science, software engineering, or a related field, or equivalent experience

Nice to have

  • Experience operating AI applications, agent platforms, or LLM services
  • Familiarity with Temporal, LangGraph, MCP, Langfuse, ClickHouse, Redis/Valkey, or S3-compatible storage
  • Experience with OpenTelemetry, Datadog APM, or Prometheus/Grafana
  • Experience with self-hosted Kubernetes, OpenShift, Kubernetes operators, CloudNativePG, GPU clusters, or AI data centers
  • Experience building AMD64 and ARM64 container images, optimizing BuildKit pipelines, or securing software supply chains

Benefits

  • Competitive salary
  • Generous benefits package

Relevance

More opportunities

Similar jobs

The newest open roles in DevOps & SRE.

Questions, answered

Frequently asked questions