Senior Application Engineer
Join a pioneer in software-defined storage over NVMe/TCP to help customers deploy, validate, and operate next-generation AI inference KV Cache solutions. This hands-on, customer-facing role involves leading technical engagements, optimizing Kubernetes-based AI inference environments, and collaborating with Product and Engineering to enhance deployment automation and documentation.
Responsibilities
- Lead technical customer engagements from discovery through production deployment.
- Design, deploy, and support Kubernetes-based AI inference environments, including GPU and KV Cache infrastructure.
- Validate AI inference workloads, troubleshoot performance bottlenecks, and optimize system configurations.
- Own customer POVs, including success criteria, test plans, issue resolution, and production handoff.
- Diagnose complex issues across Linux, Kubernetes, networking, GPU infrastructure, and application layers.
- Partner with Product and Engineering teams to drive product improvements based on customer feedback.
- Develop automation, tools, scripts, and technical documentation to improve deployment efficiency and repeatability.
- Create customer-facing technical content, including deployment guides, runbooks, best practices, and troubleshooting documentation.
Requirements
- Experience leading AI, HPC, Kubernetes, or infrastructure POCs into production environments.
- Strong Linux and Kubernetes administration, troubleshooting, and operational experience.
- Understanding of AI inference infrastructure, GPU-based deployments, and performance optimization.
- Ability to troubleshoot distributed systems using logs, metrics, dashboards, and observability tools.
- Fluency with AI tools including Claude, Co-pilot, and others.
- Excellent written and verbal communication skills with customer-facing experience.
- Ability to manage technical discussions, define success criteria, and communicate risks and tradeoffs effectively.
Nice to have
- Experience with NVIDIA Triton, vLLM, TensorRT-LLM, SGLang, KServe, Ray Serve, or similar inference platforms.
- Familiarity with GPU infrastructure, CUDA, RDMA, NCCL, and high-performance networking.
- Experience with Kubernetes Operators, Helm, and observability platforms.
- Background in Solutions Engineering, Field Engineering, SRE, Platform Engineering, or Customer Success Engineering.
- Experience creating technical blogs, tutorials, reference architectures, or customer enablement content.