Infrastructure Engineering Manager
Lead an engineering team that builds and operates infrastructure platforms supporting internal R&D teams. Own technical direction, delivery, customer support, and team development across Linux systems, automation, server fleets, and networking.
Responsibilities
- Lead and grow a team responsible for bare-metal provisioning, virtual machine infrastructure, server fleet automation, CI/CD infrastructure, customer debug support, and networking environments.
- Set the technical roadmap, priorities, execution plans, and delivery commitments for infrastructure initiatives.
- Manage virtual machine inventory and lifecycle across Linux distributions, including image readiness, OS compatibility, package baselines, kernel configuration, provisioning, and availability.
- Build and improve infrastructure capabilities for engineering and verification teams’ provisioning, testing, validation, and debugging workflows.
- Lead customer support and complex system debugging, including triage, root-cause analysis, bottleneck removal, workflow improvements, and recovery.
- Provide technical direction for Linux automation platforms, including server lifecycle management, OS installation, resource allocation, observability, and production readiness.
- Partner with firmware, driver, hardware, software, cloud, and verification teams to define requirements and deliver reliable infrastructure.
Requirements
- BSc in computer engineering, computer science, electrical engineering, or a related technical field, or equivalent experience.
- At least 8 years of experience in Linux systems administration, infrastructure automation, DevOps, system software, firmware or lab infrastructure, or a related engineering area.
- At least 3 years leading or managing engineering teams, technical projects, or cross-functional infrastructure initiatives.
- Strong Linux knowledge, including systemd, package management, kernel parameters, GRUB, sysctl, NFS, networking, boot flows, and service management.
- Hands-on experience designing, implementing, and debugging automation software using Python, scripting, and CI/CD workflows.
- Experience managing infrastructure across multiple Linux distributions, including OS images, compatibility, provisioning, package dependencies, and environment consistency.
- Ability to support internal customers through issue triage, root-cause analysis, workflow optimization, incident handling, and cross-functional communication.
- People leadership experience, including coaching, mentoring, performance management, hiring, feedback, and prioritization.
Nice to have
- Experience supporting firmware R&D, hardware bring-up, driver development, lab automation, cloud provisioning, or large-scale engineering environments.
- Knowledge of high-speed networking technologies such as RDMA, InfiniBand, Ethernet, OFED, SR-IOV, or VFIO/IOMMU.
- Experience with Ansible, infrastructure as code, Jenkins, Kubernetes, Docker, KVM, QEMU, libvirt, Vagrant, or multiple processor architectures.
- Familiarity with NVIDIA or Mellanox hardware and related firmware, diagnostics, and server recovery tools.
- Experience with fleet recovery and operations tools such as Redfish, iDRAC, iLO, IPMI, BIOS automation, and BMC configuration.
Benefits
- Competitive salary
- Benefits package