Senior Software Engineer, Profiling Services
Develop an always-on GPU profiling service that operates reliably in production, scales across clusters, and provides actionable insights for machine learning workloads. The role spans system software, drivers, CUDA, and ML workflow integration.
Responsabilidades
- Develop low-overhead, high-reliability C/C++ implementations within defined CPU and memory budgets.
- Lead end-to-end feature delivery across user-mode components, driver and platform layers, and performance counter or trace providers.
- Design profiling models that integrate with ML and AI workflows such as PyTorch and XLA.
- Correlate application events with GPU metrics to identify bottlenecks and produce actionable performance insights.
Requisitos
- Bachelor’s or master’s degree in Computer Engineering, Computer Science, or a related field, or equivalent experience.
- At least 5 years of system-level C/C++ development experience.
- Experience with concurrency, memory management, performance engineering, operating systems, computer architecture, and production software delivery.
- Strong written and verbal communication skills, with the ability to collaborate across organizations and with external partners.
Se valora
- Experience with CPU/GPU profiling and tracing stacks such as CUPTI, Nsight, performance counters, and event correlation.
- Deep knowledge of CUDA, GPU architecture, runtime and driver APIs, streams, graphs, and kernel behavior.
- Experience building continuous or multi-client profiling systems with predictable overhead at scale.
- Experience optimizing ML training or inference using profiling analysis and ecosystems such as PyTorch or JAX.
- Experience with user-mode driver development and platform security or permissions models.