Senior Solutions Architect, Cloud Infrastructure and DevOps
A hardware and infrastructure technology company is seeking a senior solutions architect to advise customers and partners on large-scale cloud, AI, and HPC infrastructure. The role combines customer consulting, system architecture, Kubernetes platforms, GPU-accelerated environments, automation, and6
Responsabilidades
- Advise on and help maintain large-scale AI and computational infrastructure.
- Troubleshoot across bare metal, operating systems, software, containers, networking, and storage.
- Assess customer environments and design production-ready Kubernetes platforms with enterprise networking and storage.
- Lead customer accounts on DevOps, platform architecture, and long-term infrastructure decisions.
- Develop and document technical methodologies, operational guidelines, runbooks, onboarding materials, and best-practice guides.
- Support development activities and lead proofs of concept and validation exercises.
- Collaborate with customers, partners, and internal teams as a trusted technical advisor.
Requisitos
- Bachelor’s, master’s, or doctoral degree in computer science, electrical or computer engineering, physics, mathematics, or a related field.
- At least 5 years of professional experience managing scalable cloud environments and working in automation engineering roles.
- Experience managing HPC or AI clusters, including deployment, optimization, and troubleshooting.
- Hands-on experience deploying and optimizing NVIDIA GPU-accelerated infrastructure, including driver management, CUDA integration, and workload profiling.
- Extensive Kubernetes experience covering orchestration, scheduling, scaling, and GPU or HPC workload integration.
- Strong knowledge of HPC and AI hardware and software, including CPUs, GPUs, and high-speed interconnects.
- Deep Linux knowledge, including Red Hat or Ubuntu, operating-system security, and relevant protocols.
- Proficiency in Python, Bash, configuration management, and infrastructure-as-code tools such as Ansible or Terraform.
- Experience with observability tools such as Grafana, Loki, and Prometheus.
- Strong solution architecture and customer consulting experience, including architectural reviews and executive presentations.
Se valora
- Knowledge of CI/CD pipelines and Kubernetes operators for GPU and network management.
- Hands-on experience with SLURM, MPI, enroot, and job provisioning.
- Experience managing software changes across cluster compute, networking, and storage.
- Experience with GPU cluster provisioning and management platforms.
- Experience with RDMA-based InfiniBand or RoCE fabrics in HPC or AI environments.