DevOps Platform Engineer
Seeking a DevOps Platform Engineer to design, build, and operate scalable infrastructure on AWS and Kubernetes. The role involves CI/CD automation, observability, and AI-driven improvements. Requires 3-5 years of experience with AWS, Kubernetes, Terraform, and scripting in Python/Go/Bash.
Responsibilities
- Design, build, and maintain application infrastructure and platforms.
- Develop scalable, reliable, and secure infrastructure on AWS.
- Build and improve CI/CD pipelines and infrastructure automation.
- Manage and operate Kubernetes-based production environments.
- Develop self-service tools to improve developer productivity.
- Improve system reliability, availability, and scalability.
- Define and implement best practices for infrastructure, deployments, observability, and security.
- Lead troubleshooting and root cause analysis for complex production issues.
- Collaborate with cross-functional teams including R&D, Security, and Cloud.
- Evaluate and implement AI-based tools for automation and operational efficiency.
- Develop AI agents and LLM-based tools for platform engineering.
Requirements
- 3-5 years of experience in DevOps, Platform Engineering, or Infrastructure Engineering.
- Strong hands-on experience with AWS in production.
- Proficiency with AWS services such as EKS, EC2, IAM, VPC, Route 53, S3, RDS, Lambda, CloudWatch.
- Strong experience with Kubernetes and containers.
- Hands-on experience with Terraform.
- Experience building CI/CD pipelines using GitHub Actions, GitLab CI, Argo CD, or similar.
- Proficiency in scripting with Python, Go, or Bash.
- Strong understanding of Linux, networking, DNS, load balancing, and distributed systems.
- Experience with monitoring, logging, tracing, and observability platforms.
- Strong troubleshooting and incident investigation skills.
- Autonomous and ownership-driven work style.
- Strong English communication and collaboration skills.
- Hands-on experience using AI and Generative AI tools in development, infrastructure, or operations.
- Experience with LLM APIs, AI coding assistants, or AI automation tools.
- Ability to identify processes for AI improvement and implement solutions.
- Understanding of prompt engineering, responsible AI, and secure handling of information.
- Advantage: experience with CAG, AI agents, RAG, MCP, LLM integrations, or AIOps.
Nice to have
- Experience with data platforms or large-scale data processing systems.
- Understanding of data pipelines, streaming systems, or distributed data architectures.
- Experience with multi-region or multi-account AWS environments.
- Experience with GitOps tools like Argo CD or Akuity.
- Experience with service mesh, API gateways, or service discovery.
- Experience with Datadog, Prometheus, Grafana, OpenTelemetry.
- Experience with cloud security, secrets management, and Vault.