Senior LLM Agents Architect
Build production-grade agentic AI systems that optimize GPU compute kernels, analyze simulation and profiling data, and accelerate GPU hardware-software co-design.
תחומי אחריות
- Design and build agentic AI systems for generating, analyzing, and optimizing GPU compute kernels.
- Encode GPU architecture and performance expertise into agent workflows and optimization systems.
- Develop performance-forensics agents that analyze simulation traces and profiler data to identify bottlenecks and recommend mitigations.
- Create agentic workflows for GPU architectural studies and what-if analysis across hardware configurations.
- Explore LLM-driven approaches to compiler optimization, code generation, and hardware-software co-design.
- Prototype, productize, and integrate agent systems with internal services and GPU capabilities.
- Establish offline evaluation sets, online telemetry, cost controls, safety checks, and reliable iteration processes.
- Mentor teams on agent orchestration, prompting, RAG, observability, documentation, and operational playbooks.
דרישות
- At least 8 years of experience in applied ML/AI or large-scale systems, including at least 2 years building production agentic or LLM-powered applications.
- Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field.
- Strong knowledge of computer architecture, including memory hierarchies, parallelism, pipelining, cache behavior, and NVIDIA GPU architecture.
- Hands-on CUDA experience writing, profiling, and optimizing GPU kernels, with proficiency in Nsight Compute, Nsight Systems, or equivalent tools.
- Ownership of an end-to-end production agentic or LLM application, from requirements and architecture through evaluation and hardening.
- Strong Python and systems programming skills, preferably in C++.
- Experience with tool use, RAG pipelines, model adaptation, AI observability, evaluation, telemetry, guardrails, and rollback plans.
- Ability to translate hardware and software domain expertise into tools, constraints, and evaluation metrics.
- Strong communication, documentation, facilitation, and cross-functional collaboration skills.
יתרון
- Experience with PyTorch compilation, TorchDynamo, TorchInductor, Triton, PTX, graph compilers, kernel fusion, or auto-tuning.
- Background in HPC or GPU performance engineering, performance modeling, or hardware simulation.
- Experience with distributed processing, multi-GPU workloads, NVLink, or InfiniBand.
- Familiarity with coding-agent tools and frameworks, including tool orchestration, context management, evaluation, and failure recovery.
- Experience building domain-specific coding agents.
הטבות
- Competitive salary and comprehensive benefits package for employees and families.