Senior AI Engineer, LLM Agent Systems
Develop production-grade LLM agent systems that optimize GPU kernels, analyze performance data, support hardware architecture studies, and improve AI-driven software engineering workflows. The role partners closely with hardware architects, verification engineers, performance specialists, and other
Responsibilities
- Design and build agentic AI systems that generate, analyze, and optimize GPU compute kernels
- Encode GPU architecture and performance expertise into agent workflows
- Build performance-forensics agents using simulation traces and profiler data to identify bottlenecks and recommend mitigations
- Develop agentic workflows for GPU architectural studies and configuration analysis
- Explore LLM-driven hardware/software co-design, optimization, and code-generation pipelines
- Prototype, productize, and integrate agent systems with internal services and GPU capabilities
- Establish offline evaluation sets, online telemetry, cost controls, and safety mechanisms
- Mentor teams on agent orchestration, prompting, retrieval, observability, documentation, and operational playbooks
Requirements
- 8+ years of experience in applied machine learning, artificial intelligence, or large-scale systems, including 2+ years building agentic or LLM-powered production applications
- Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field
- Strong knowledge of computer architecture, including memory hierarchies, parallelism, pipelining, and cache behavior
- Essential familiarity with NVIDIA GPU architecture, including streaming multiprocessors, warp scheduling, memory models, and occupancy
- Hands-on CUDA kernel development, profiling, and optimization experience
- Experience using Nsight Compute, Nsight Systems, or comparable profiling tools
- Ownership of an end-to-end production agentic or LLM application, from requirements and architecture through evaluation and hardening
- Strong Python skills and proficiency in a systems language, preferably C++
- Experience with tool use, retrieval-augmented generation, and model adaptation techniques
- Ability to translate hardware and software experts' heuristics into tools, constraints, and evaluation metrics
- Experience building AI observability, including dataset and version management, offline testing, telemetry, guardrails, and rollback plans
- Strong communication, documentation, facilitation, and cross-functional collaboration skills
Nice to have
- Experience with PyTorch compilation and lowering, TorchDynamo, TorchInductor, Triton, PTX, graph compilers, kernel fusion, or auto-tuning
- HPC or GPU performance engineering experience, including performance modeling or hardware simulators
- Experience with distributed processing, multi-GPU workloads, NVLink, or InfiniBand
- Familiarity with frontier coding agents and their orchestration, context management, and autonomous execution patterns
- Experience building domain-specific coding agents with agent harnesses or frameworks such as LangChain, LangGraph, or CrewAI
Benefits
- Competitive salary
- Comprehensive benefits package for employees and families