Sustaining Lead Engineer
Own the manufacturing health of server products in volume production, improving producibility, yield, quality, reliability, and supply continuity across manufacturing sites.
Responsibilities
- Define production-readiness criteria and validate manufacturing processes during new product introduction and transition to volume production
- Own yield, quality, and reliability metrics for server platforms in volume production; investigate root causes and drive corrective actions
- Work with contract manufacturers to monitor process health, resolve production escapes, and improve high-volume production lines
- Lead debug and failure analysis for field returns and fleet reliability events; establish containment and drive permanent fixes
- Qualify second-source electrical components with supply chain partners to reduce supply risk and maintain continuity
- Develop sustaining test strategies, failure-analysis workflows, screening methods, and data-based disposition criteria
- Analyze manufacturing and field data to identify trends and systemic failure modes, then drive design or process changes
- Travel to manufacturer sites four to six times per year for week-long trips
Requirements
- At least 8 years developing and supporting hardware products
- Experience leading cross-functional efforts across hardware, firmware, test, and manufacturing teams
- Knowledge of server or complex electronic system architecture, including PCBs, power delivery, and high-speed interfaces such as DDR4/5 and PCIe Gen3/4/5
- Hands-on experience developing manufacturing tests, debugging, and analyzing data at scale
- Practical scripting experience with Python or Bash
Nice to have
- Experience with contract manufacturers or ODMs in sustaining or production engineering
- Experience with volume manufacturing, yield management, failure analysis, root-cause investigation, and corrective or preventive actions
- System-level knowledge of networking, storage, or Linux
- Technical leadership in a matrixed, multi-site environment