Senior Cloud and AI Platform Architect
Own the architecture and evolution of a large-scale GCP platform and its production agentic AI ecosystem. Lead cloud reliability, delivery standards, security, cost efficiency, and AI governance for a real-time sports data company.
Responsabilidades
- Own GKE fleet design, VPC networking, IAM, Workload Identity, and multi-environment cloud strategies across GCP and relevant AWS environments
- Lead Infrastructure as Code standards, GitOps deployment patterns, and secure CI/CD paths for R&D teams
- Define reliability architecture, disaster recovery strategy, observability platforms, incident reduction practices, and MTTR improvements
- Drive FinOps through rightsizing, idle-resource elimination, spend attribution, cost-aware design reviews, and cloud cost optimization
- Design the agentic AI ecosystem, including orchestration, tool and MCP interfaces, vector databases, memory strategies, and production deployment patterns
- Build AI evaluation and guardrail capabilities with permission models, human review checkpoints, audit logging, and policy-as-code
- Own AI cost observability, token monitoring, model routing, and alerts for excessive model usage
- Define reference architectures, RAG patterns, prompt management, and model API standards for AI feature teams
- Establish platform and AI architecture standards adopted across engineering
Requisitos
- 8+ years of experience in platform or infrastructure engineering, DevOps, or systems architecture
- Proven ownership of enterprise-scale platform architecture decisions
- Deep production expertise with Kubernetes, preferably GKE, container runtimes, and service mesh technologies
- Advanced Infrastructure as Code experience with Terraform or Pulumi and GitOps experience with ArgoCD
- Strong GCP architecture expertise across networking, IAM, Workload Identity, and cloud cost management
- At least 1–2 years of hands-on experience delivering LLM-based or agentic systems to production
- Experience with agent frameworks, tool or function calling, and production Retrieval-Augmented Generation systems
- Demonstrated reliability and cost outcomes, such as uptime improvement, MTTR reduction, or cloud spend optimization
- Experience with Apache Kafka or comparable large-scale streaming platforms
- Ability to lead through influence, mentor engineers, and communicate architectural trade-offs to technical and executive audiences
Se valora
- Experience with Model Context Protocol, multi-agent orchestration, or automated AI evaluation frameworks
- FinOps certification or significant cloud cost-reduction experience
- Production experience with Vertex AI or AWS Bedrock
- Experience supporting GPU and inference workloads
- Deep AWS architecture experience alongside GCP
Beneficios
- Hybrid work model
- Work from anywhere for up to 30 days per year
- Extended parental leave
- Quarterly company-wide recharge days
- Health insurance
- Birthday day off
- Cibus
- Holiday eves off
- Team and company social events
- Dog-friendly office
- Kids summer camp
- Access to a wellbeing platform