Senior Software Architect, AI Networking
Hands-on architect shaping LLM inference at scale. Design scalable multi-node GPU cluster architectures, optimize latency, throughput, and cost, prototype KV cache and parallelism approaches, and integrate networking technologies. Collaborate across model, systems, compiler, and networking teams. Requires 8+ years distributed systems experience, deep learning systems, C++/Python, CUDA.
Responsibilities
- Design and evolve scalable architectures for multi-node LLM inference across GPU clusters.
- Develop infrastructure to optimize latency, throughput, and cost efficiency for serving large models.
- Collaborate with model, systems, compiler, and networking teams for high-performance solutions.
- Prototype approaches for KV cache handling, tensor/pipeline parallelism, and dynamic batching.
- Evaluate and integrate software/hardware technologies like load balancing, telemetry, congestion control.
- Work with internal and external partners to translate architecture into reliable systems.
- Author design docs, specs, and technical blog posts; contribute to open source.
Requirements
- Bachelor’s, Master’s, or PhD in CS/EE or equivalent experience.
- 8+ years in large-scale distributed systems or performance-critical software.
- Deep understanding of deep learning systems, GPU acceleration, AI model execution flows, or high-performance networking.
- Solid software engineering in C++ and/or Python, preferably CUDA.
- Strong system-level thinking across memory, networking, scheduling, compute orchestration.
- Excellent communication and cross-domain collaboration.
Nice to have
- Experience with LLM training/inference pipelines, transformer optimization, model-parallel deployments.
- Proven performance profiling and optimization across LLM stack.
- AI accelerators and distributed communication patterns, congestion control, load balancing.
- Optimization process for complex systems deployed at scale.
- Passion for tough technical problems and high-impact solutions.