← Back to jobs
N

Senior Technical Support Engineer

NVIDIA·Israel·en
Not specifiedFull-timeCustomer SupportSemiconductorsComputer Hardware

Support customers running production Slurm clusters for AI and high-performance computing. Own complex cases from investigation through resolution, diagnose issues across Slurm and related infrastructure, and help customers operate reliable, scalable clusters.

Responsibilities

  • Own customer support cases for production AI and HPC clusters from investigation through resolution.
  • Diagnose issues involving Slurm daemons, scheduling, node management, resource allocation, accounting, authentication, and high availability.
  • Resolve configuration and policy issues involving partitions, reservations, priorities, fair share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints.
  • Investigate performance, reliability, and scalability problems using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging.
  • Isolate issues across Slurm and related Linux, MUNGE, database, networking, parallel storage, container, GPU, and cluster-management systems.
  • Advise customers on configuration, upgrades, operational practices, resource management, and production incident recovery.
  • Collaborate with engineering teams by preparing technical findings, reproducible test cases, and defect reports.
  • Create guides, knowledge-base articles, diagnostic tools, and internal training materials.

Requirements

  • At least 5 years administering and supporting Slurm in production HPC or AI environments, including business-critical incidents.
  • Expert understanding of Slurm architecture, daemons, configuration, scheduling, accounting, resource management, and failure modes.
  • Ability to independently investigate complex Slurm incidents and guide them to resolution.
  • In-depth Linux system administration experience, including systemd, cgroups, authentication, networking, and database-backed services.
  • Experience operating multi-user Slurm clusters with complex scheduling policies and heterogeneous compute resources.
  • Strong analytical and research skills to distinguish Slurm defects from configuration, integration, infrastructure, and workload issues.
  • Excellent written and verbal communication skills, including explaining technical findings and recommendations.
  • Bachelor’s degree in computer science, engineering, or a related field, or equivalent experience.

Nice to have

  • Experience supporting Slurm clusters with thousands of nodes or GPUs.
  • Experience diagnosing scheduler performance, job throughput, controller load, and database scaling.
  • Familiarity with Slurm source code, plugins, SPANK, Lua job-submit plugins, or upstream issue investigation.
  • Experience with Pyxis, Enroot, Apptainer, Singularity, or similar HPC container technologies.
  • Experience integrating Slurm with NVIDIA Base Command Manager, Bright Cluster Manager, or another cluster-management platform.

Relevance

More opportunities

Similar jobs

The newest open roles in Customer Support.

Questions, answered

Frequently asked questions