← Back to jobs
N

Senior LLM Agents Architect

NVIDIA·Israel·en
On-siteFull-timeAI EngineeringSemiconductors

Build production-grade agentic AI systems that optimize GPU compute kernels, analyze simulation and profiling data, and accelerate GPU hardware-software co-design.

Responsibilities

  • Design and build agentic AI systems for generating, analyzing, and optimizing GPU compute kernels.
  • Encode GPU architecture and performance expertise into agent workflows and optimization systems.
  • Develop performance-forensics agents that analyze simulation traces and profiler data to identify bottlenecks and recommend mitigations.
  • Create agentic workflows for GPU architectural studies and what-if analysis across hardware configurations.
  • Explore LLM-driven approaches to compiler optimization, code generation, and hardware-software co-design.
  • Prototype, productize, and integrate agent systems with internal services and GPU capabilities.
  • Establish offline evaluation sets, online telemetry, cost controls, safety checks, and reliable iteration processes.
  • Mentor teams on agent orchestration, prompting, RAG, observability, documentation, and operational playbooks.

Requirements

  • At least 8 years of experience in applied ML/AI or large-scale systems, including at least 2 years building production agentic or LLM-powered applications.
  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field.
  • Strong knowledge of computer architecture, including memory hierarchies, parallelism, pipelining, cache behavior, and NVIDIA GPU architecture.
  • Hands-on CUDA experience writing, profiling, and optimizing GPU kernels, with proficiency in Nsight Compute, Nsight Systems, or equivalent tools.
  • Ownership of an end-to-end production agentic or LLM application, from requirements and architecture through evaluation and hardening.
  • Strong Python and systems programming skills, preferably in C++.
  • Experience with tool use, RAG pipelines, model adaptation, AI observability, evaluation, telemetry, guardrails, and rollback plans.
  • Ability to translate hardware and software domain expertise into tools, constraints, and evaluation metrics.
  • Strong communication, documentation, facilitation, and cross-functional collaboration skills.

Nice to have

  • Experience with PyTorch compilation, TorchDynamo, TorchInductor, Triton, PTX, graph compilers, kernel fusion, or auto-tuning.
  • Background in HPC or GPU performance engineering, performance modeling, or hardware simulation.
  • Experience with distributed processing, multi-GPU workloads, NVLink, or InfiniBand.
  • Familiarity with coding-agent tools and frameworks, including tool orchestration, context management, evaluation, and failure recovery.
  • Experience building domain-specific coding agents.

Benefits

  • Competitive salary and comprehensive benefits package for employees and families.

Relevance

More opportunities

Similar jobs

The newest open roles in AI Engineering.

Questions, answered

Frequently asked questions