Senior Software Engineer - AI Agent Platform
Own and evolve the production runtime powering AI agents, focusing on reliability, scalability, and observability. Build platform capabilities for LLM orchestration, tool calling, and state management, collaborating across teams to deliver new AI capabilities safely and quickly.
Responsibilities
- Own and evolve the production runtime for executing and orchestrating AI agents.
- Design platform capabilities for agent execution, tool calling, streaming, state management, persistence, and long-running workflows.
- Build resilient integrations with multiple LLM providers and model-serving platforms.
- Design provider-routing and fallback strategies based on availability, latency, quality, and cost.
- Implement retries, timeouts, circuit breakers, rate-limit handling, idempotency, and graceful degradation.
- Ensure platform availability during outages of external dependencies or infrastructure.
- Build reliable mechanisms for loading, caching, versioning, and recovering agent configurations and artifacts.
- Create end-to-end observability for AI requests, including model, provider, agent, latency, token usage, cost, errors, retries, and fallbacks.
- Define dashboards, alerts, SLOs, and runbooks for production AI workloads.
- Lead investigation of complex production issues across application code, infrastructure, external providers, distributed state, and agent behavior.
- Improve platform scalability, concurrency, latency, and resource efficiency.
- Build reusable APIs and abstractions to help agent developers add capabilities without duplicating infrastructure logic.
- Strengthen platform quality through integration testing, load testing, failure injection, and dependency-outage simulations.
- Turn production incidents into architectural improvements, automated tests, monitoring, and operational safeguards.
- Collaborate with product, infrastructure, and engineering teams to translate customer and business requirements into platform capabilities.
- Mentor engineers and establish best practices for building and operating reliable production AI systems.
Requirements
- 7+ years of professional software engineering experience, primarily in backend, platform, or distributed systems.
- Expert-level TypeScript and Node.js skills.
- Experience with NestJS or a comparable backend framework.
- Proven experience designing, building, and operating large production services.
- Strong understanding of distributed-systems patterns: retries, backoff, idempotency, circuit breakers, caching, consistency, failure recovery.
- Systematic debugging mindset with ability to trace failures across multiple services and dependencies.
- Experience owning customer-facing systems where availability, latency, and correctness directly affect users.
- Strong experience with cloud infrastructure and managed services, preferably AWS.
- Experience with distributed caching and storage technologies such as Redis and S3.
- Hands-on experience with production observability: structured logs, metrics, tracing, dashboards, alerts, SLOs.
- Experience participating in incident response and driving follow-up improvements.
- Strong API design, testing, and software architecture fundamentals.
- Excellent communication and collaboration skills across engineering, product, infrastructure, and AI teams.
- Bachelor’s degree in Computer Science or a related field, or equivalent practical experience.
- Hands-on experience with or strong understanding of LLM APIs, streaming, tool calling, and structured outputs.
- Understanding of agentic patterns such as reasoning loops, tool orchestration, and multi-step workflows.
- Knowledge of tokens, context windows, rate limits, latency, and model-specific behavior.
- Experience with multi-provider LLM integrations and trade-offs between providers and models.
- Understanding how retries and fallbacks affect response quality, latency, correctness, and cost.
- Knowledge of observability required to understand a request across multiple LLM and tool calls.
- Familiarity with techniques for controlling and optimizing LLM usage and cost.
Nice to have
- Experience building an AI gateway, agent runtime, inference platform, or workflow engine.
- Experience with OpenAI, Gemini/Vertex AI, AWS Bedrock, or similar platforms.
- Familiarity with model and provider routing based on quality, availability, latency, and cost.
- Experience with Kafka or other event-driven architectures.
- Familiarity with MCP or agent-to-agent communication protocols.
- Experience with Kubernetes and cloud-native infrastructure.
- Experience with Grafana or similar tooling.
- Experience with chaos engineering, fault injection, or large-scale load testing.
- Experience building internal developer platforms or frameworks used by multiple engineering teams.
- Familiarity with conversational AI, RAG, memory systems, or generative user experiences.
Benefits
- Opportunity to shape a production AI platform used by customers worldwide.
- Complex engineering challenges at the intersection of distributed systems and generative AI.
- Direct influence over platform architecture, reliability, and technical direction.
- Collaborative environment with experienced engineers across AI, product, backend, and infrastructure.
- Opportunities for professional growth, mentorship, and technical leadership.