Senior Software Engineer - AIOps Platform
Architect and build distributed systems for high-scale telemetry ingestion and AI model serving. Develop agentic AIOps for GPU cluster monitoring and incident response. Use Go, C++, Rust, and Kubernetes.
Responsibilities
- Architect and build agentic AIOps system for GPU fleet health monitoring and incident response
- Design distributed systems for high-scale telemetry ingestion and real-time analysis
- Research and prototype data storage strategies across diverse database technologies
- Build and own model-serving infrastructure for SaaS and on-premises deployments
- Instrument services with deep observability for rapid debugging and performance improvement
- Contribute to core libraries and abstractions for the AIOps engineering team
Requirements
- B.Sc./M.Sc. in Computer Science, Computer Engineering, or related field
- 5+ years of software engineering experience building production distributed systems
- Expert-level proficiency in Go, C++, or Rust with focus on high-performance concurrent architectures
- Solid understanding of Kubernetes and container-based deployments
- Experience deploying, monitoring, and maintaining ML models or data-intensive services
- Comfort working in ambiguous, fast-moving environments
Nice to have
- Experience building ML model-serving platforms or MLOps tooling at scale
- Track record of taking systems from prototype to production-grade platform
- Systems thinker understanding full stack from data transfer to distributed cluster processing
- Ability to simplify complex problems and build internal tools for other engineering teams
Benefits
- Competitive salaries
- Generous benefits package