Эта вакансия пока только на английском.

← Назад к вакансиям
N

Senior LLM Agents Architect

NVIDIA·Израиль·en
В офисеПолная занятостьAI EngineeringSemiconductors

Build production-grade agentic AI systems that optimize GPU compute kernels, analyze simulation and profiling data, and accelerate GPU hardware-software co-design.

Обязанности

  • Design and build agentic AI systems for generating, analyzing, and optimizing GPU compute kernels.
  • Encode GPU architecture and performance expertise into agent workflows and optimization systems.
  • Develop performance-forensics agents that analyze simulation traces and profiler data to identify bottlenecks and recommend mitigations.
  • Create agentic workflows for GPU architectural studies and what-if analysis across hardware configurations.
  • Explore LLM-driven approaches to compiler optimization, code generation, and hardware-software co-design.
  • Prototype, productize, and integrate agent systems with internal services and GPU capabilities.
  • Establish offline evaluation sets, online telemetry, cost controls, safety checks, and reliable iteration processes.
  • Mentor teams on agent orchestration, prompting, RAG, observability, documentation, and operational playbooks.

Требования

  • At least 8 years of experience in applied ML/AI or large-scale systems, including at least 2 years building production agentic or LLM-powered applications.
  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field.
  • Strong knowledge of computer architecture, including memory hierarchies, parallelism, pipelining, cache behavior, and NVIDIA GPU architecture.
  • Hands-on CUDA experience writing, profiling, and optimizing GPU kernels, with proficiency in Nsight Compute, Nsight Systems, or equivalent tools.
  • Ownership of an end-to-end production agentic or LLM application, from requirements and architecture through evaluation and hardening.
  • Strong Python and systems programming skills, preferably in C++.
  • Experience with tool use, RAG pipelines, model adaptation, AI observability, evaluation, telemetry, guardrails, and rollback plans.
  • Ability to translate hardware and software domain expertise into tools, constraints, and evaluation metrics.
  • Strong communication, documentation, facilitation, and cross-functional collaboration skills.

Будет плюсом

  • Experience with PyTorch compilation, TorchDynamo, TorchInductor, Triton, PTX, graph compilers, kernel fusion, or auto-tuning.
  • Background in HPC or GPU performance engineering, performance modeling, or hardware simulation.
  • Experience with distributed processing, multi-GPU workloads, NVLink, or InfiniBand.
  • Familiarity with coding-agent tools and frameworks, including tool orchestration, context management, evaluation, and failure recovery.
  • Experience building domain-specific coding agents.

Условия и преимущества

  • Competitive salary and comprehensive benefits package for employees and families.

Соответствие

Больше возможностей

Похожие вакансии

Новые вакансии в категории «AI Engineering».

Создайте профиль, чтобы увидеть оценку соответствия.

Создайте профиль, чтобы увидеть оценку соответствия.

Создайте профиль, чтобы увидеть оценку соответствия.

Создайте профиль, чтобы увидеть оценку соответствия.

Senior LLM Agent Architect for GPU Optimization Systems | CVZilla