Senior Software Engineer
Senior Software Engineer responsible for building and operating a high-scale AIOps platform that ingests telemetry from GPU clusters, correlates events, and orchestrates diagnostic workflows. Work involves distributed systems, data storage strategies, observability instrumentation, and deploying ML models to production in SaaS and on-premises environments.
Responsibilities
- Architect and build an agentic AIOps system for GPU fleet health monitoring, telemetry correlation, alerting, and automated diagnostics.
- Research and evaluate data storage strategies and representations for AI model training and predictive accuracy.
- Design distributed systems to handle extreme telemetry density from large-scale AI clusters.
- Instrument services with deep observability (metrics, logs, traces) for debugging and performance improvement.
- Build and own model-serving infrastructure, packaging, versioning, deploying, and monitoring AI models.
- Contribute to core libraries and abstractions to accelerate development across the team.
Requirements
- B.Sc./M.Sc. in Computer Science, Computer Engineering, or related technical field.
- 5+ years of software engineering experience building production distributed systems.
- Expert-level proficiency in Go, C++, or Rust, with focus on high-performance, concurrent architectures.
- Solid understanding of Kubernetes and container-based deployments.
- Experience deploying, monitoring, and maintaining ML models or data-intensive services in production.
- Comfort working in ambiguous, fast-moving environments.
Nice to have
- Experience building ML model-serving platforms or MLOps tooling at scale.
- Track record of taking systems from prototype to production-grade platform.
- Systems thinker with full-stack understanding of data flow and distributed processing.
- Ability to simplify complex problems and build internal tools/frameworks.
Benefits
- Competitive salaries
- Generous benefits package