Senior DevOps Engineer
Own reliable, secure, and automated cloud operations for a high-traffic enterprise customer experience platform. Improve delivery speed, observability, scalability, and operational quality through modern DevOps practices and AI-enabled tooling.
Responsibilities
- Build and maintain GitHub Actions CI/CD pipelines for automated testing, linting, security checks, and frequent deployments
- Operate GitOps workflows with ArgoCD as the source of truth for cluster state
- Manage AWS environments across multiple accounts using infrastructure as code
- Run and troubleshoot Kubernetes EKS workloads and maintain shared Helm charts
- Improve edge performance, security, reliability, observability, and developer experience
- Investigate production incidents, perform root-cause analysis, and address performance and scale issues
- Use AI assistants, AIOps, and automation to accelerate delivery and reduce repetitive operational work
- Collaborate with R&D to improve service performance, deployment, failover, reliability, and scalability
Requirements
- At least 4 years of experience as a DevOps, SRE, or platform engineer in high-scale, high-traffic environments
- At least 4 years of experience with AWS, including EC2, RDS, VPC, S3, Lambda, CloudFront, and IAM
- Hands-on experience with Docker and container orchestration, including Kubernetes or EKS and Helm
- Experience building CI/CD pipelines and GitOps workflows using GitHub Actions, ArgoCD, or equivalent tools
- Proficiency with Bash or Python scripting and experience using AI coding assistants for tooling and troubleshooting
- Practical experience with Linux and Windows systems, system performance, and security
- Experience with GitHub, infrastructure automation, and infrastructure-as-code tools
- Experience managing and troubleshooting cloud and hybrid web environments
- Experience with MySQL, Redis, and Elasticsearch or OpenSearch
- Demonstrated interest in adopting AI tools to improve DevOps work
Nice to have
- AIOps or AI/ML experience for failure prediction, root-cause analysis, or resource optimization
- Experience with Kubernetes, CloudFormation, Terraform, and New Relic
- Experience with Cloudflare WAF, Workers, or CDN
- Experience with ELK, IIS, or directory services
- Familiarity with SAST, security scanning, and shift-left security practices
- Knowledge of infrastructure-as-code and DevOps culture