Senior DevOps/SRE Engineer
A technology company is seeking a Senior DevOps/SRE Engineer to support a SaaS/IaaS digital twin platform for AI data center infrastructure. The role focuses on automation, cloud operations, infrastructure reliability, and secure production environments.
Responsibilities
- Own DevOps, infrastructure, and Site Reliability Engineering activities for the platform.
- Automate repetitive workflows and improve operational efficiency.
- Support microservices-based architecture and cloud-scale infrastructure.
- Deploy and troubleshoot non-disruptive cloud operations in secure production environments.
- Continuously evaluate systems and drive reliability and efficiency improvements.
- Manage operating system, Kubernetes cluster, and orchestration-tool deployments and upgrades.
- Support engineering teams with CI/CD tools including Git and Jenkins.
- Manage multiple operational workstreams as priorities evolve.
Requirements
- Bachelor’s degree in Computer Science or equivalent experience.
- At least 5 years of experience with complex microservices architectures.
- Advanced Kubernetes and Docker skills.
- Experience deploying, configuring, and administering Linux-based bare-metal servers in IaaS environments.
- Strong networking knowledge, including VLANs, routing, and VPNs.
- Experience with relational databases such as MySQL and SQL.
- Experience with blue-green and canary deployment strategies for non-disruptive cloud operations.
- Infrastructure-as-code experience with tools such as Ansible and Terraform.
- Expertise in AWS.
- Knowledge of best practices for managing and monitoring highly available, secure production infrastructure.
Nice to have
- Strong Infrastructure as a Service expertise.
- Linux/Unix administration experience.
- Experience with Prometheus and Grafana.
- Experience with application performance monitoring tools such as Dynatrace, Datadog, AppDynamics, or New Relic.
- Experience implementing metrics collection and alerting infrastructure.