Technical Lead
Define technical architecture for a distributed testing and reliability platform. Build systems for large-scale workload simulation, chaos injection, and observability. Lead engineering excellence, mentor engineers, and drive quality as a quantitative discipline. Hands-on role requiring strong Python skills, deep distributed systems knowledge, and experience with storage/networking systems.
Responsibilities
- Define and own the technical architecture of the distributed testing and reliability platform.
- Lead engineering teams, setting standards and mentoring engineers.
- Build systems for massive-scale workload simulation and adversarial failure injection.
- Advance AI-driven test automation (intelligent scenario generation, LLM-augmented root cause analysis).
- Drive observability and reliability engineering, tracking latency and system health.
- Collaborate with Core R&D, Storage Kernel, and Infrastructure teams on reliability strategies.
- Establish engineering practices including design reviews and code quality standards.
Requirements
- 6+ years hands-on Python development experience.
- Deep understanding of distributed systems (concurrency, consistency, fault tolerance).
- Background in storage systems, networking (TCP/IP, RDMA), cloud infrastructure, or high-performance backends.
- Experience building large-scale infrastructure platforms or reliability engineering systems.
- Proven leadership of complex technical initiatives.
- Ability to mentor engineers and drive technical alignment across teams.
Nice to have
- Experience with storage systems, file systems, or high-performance distributed environments.
- Background in chaos engineering, fault injection, or simulation systems.
- Familiarity with observability tooling and performance engineering at scale.
- Experience building testing or reliability platforms as engineering products.
- Prior experience as a Team Lead in a high-growth infrastructure company.