Эта вакансия пока только на английском.

← Назад к вакансиям
N

Senior Software Architect, AI Inference Infrastructure

NVIDIA AI·Израиль·en
Не указаноПолная занятостьSoftware ArchitectureSemiconductors

Lead the architecture and optimization of large-scale LLM inference systems running across GPU clusters. The role spans software, hardware, networking, memory orchestration, scheduling, and production deployment, with collaboration across engineering, research, and external partner teams.

Обязанности

  • Design and evolve scalable architectures for multi-node LLM inference across GPU clusters
  • Develop infrastructure that improves latency, throughput, and cost efficiency for production model serving
  • Collaborate with model, systems, compiler, and networking teams on integrated high-performance solutions
  • Prototype KV-cache handling, tensor and pipeline parallelism, and dynamic batching approaches
  • Evaluate and integrate technologies for load balancing, telemetry, congestion control, and application integration
  • Translate high-level architecture into reliable, high-performance systems with internal teams and external partners
  • Write design documents, technical specifications, and technical blog posts, and contribute to open-source projects when appropriate

Требования

  • Bachelor’s, master’s, or doctoral degree in computer science, electrical engineering, or equivalent experience
  • At least 8 years of experience building large-scale distributed systems or performance-critical software
  • Deep understanding of deep learning systems, GPU acceleration, AI model execution flows, or high-performance networking
  • Strong software engineering skills in C++ and/or Python
  • Strong system-level understanding of memory, networking, scheduling, and compute orchestration
  • Ability to collaborate across diverse technical domains

Будет плюсом

  • Experience with LLM training or inference pipelines, transformer optimization, or model-parallel deployments
  • Experience profiling and optimizing bottlenecks across AI training or inference stacks
  • Knowledge of AI accelerators, distributed communication, congestion control, or load balancing
  • Experience optimizing complex systems deployed at scale
  • Experience driving complex organizational processes from planning through implementation

Условия и преимущества

  • Competitive salary
  • Comprehensive benefits package for employees and families

Соответствие

Больше возможностей

Похожие вакансии

Новые вакансии в категории «Software Architecture».

Создайте профиль, чтобы увидеть оценку соответствия.

Создайте профиль, чтобы увидеть оценку соответствия.

Создайте профиль, чтобы увидеть оценку соответствия.

Создайте профиль, чтобы увидеть оценку соответствия.