DevOps Engineer
Join a global DevOps and Platform team responsible for designing, operating, and modernizing highly available distributed production systems. Own initiatives from design through rollout while collaborating with platform, development, architecture, AI/ML, and product teams.
Responsabilidades
- Lead platform initiatives from design through implementation and production rollout
- Design, build, and operate highly available distributed systems on AWS and GCP
- Manage Kubernetes, Docker, Helm, CI/CD automation, data, caching, and messaging layers
- Modernize and consolidate technology stacks and architectures into a coherent platform
- Build and operate infrastructure for AI workloads, including GPU inference platforms
- Drive system reliability, monitoring, and observability
- Troubleshoot complex issues across application, infrastructure, networking, operating system, and data layers
- Collaborate with Platform, Development, Architecture, AI/ML Research, and Product teams
Requisitos
- At least 2 years of hands-on experience in DevOps, Platform, or SRE roles operating distributed production systems at scale
- Production expertise with AWS services, including EKS, ECS, S3, EC2, CloudFront, IAM, Bedrock, MSK, Lambda, and CloudWatch
- Infrastructure as Code experience with Terraform or a similar tool
- Strong hands-on experience with Kubernetes and Helm, including Kubernetes internals and production operations
- Experience with Docker and containerized systems
- Strong knowledge of networking and operating systems, including TCP/IP, DNS, load balancing, and Linux internals
- Strong experience with CI/CD and automation, including GitHub Actions
- Experience with monitoring and observability using metrics, tracing, logging, Prometheus, Grafana, or similar tools
- High proficiency with current AI coding tools and their practical application
Se valora
- Experience integrating or migrating systems across different technology stacks or organizations
- Experience leading cross-team technical initiatives
- Experience working with GPUs for AI/ML inference and serving
Beneficios
- Hybrid and flexible work environment
- Extended private health insurance, including mental health coverage
- Personal and professional development programs
- Occasional cross-company long weekends