Data Center Hardware Quality & Reliability Engineer
OpenAI · San Francisco
About this role
**Data Center Hardware Quality & Reliability Engineer** **About The Role** Own the end-to-end data-center hardware quality and reliability loop for OpenAI's 3P infrastructure and 1P current and next-gen platforms. Turn field failures into quantified risk, fast containment, verified root cause, improved MQE/NPI and manufacturing-test coverage, accurate spares forecasts, and upstream changes that prevent recurrence. **Key Responsibilities** • Build and govern the field-quality data model across telemetry, tickets, RMA/repair, FA, firmware, configuration, supplier, and manufacturing genealogy • Define AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF, repair-cycle-time, and forecast-versus-actual metrics • Provide fleet-level macro views and unit/FRU/cohort-level micro views; detect shifts and bound affected populations • Lead systemic field-failure triage, containment, failure analysis, 8D/CAPA, risk assessment, and corrective-action verification • Develop cohort, life-data, reliability-growth, and spare-demand projections by product, FRU, supplier, configuration, geography, and age • Partner with MQE and NPI to convert field mechanisms into manufacturing-test coverage and qualification gates • Define supplier/CM FA standards, field-data contracts, scorecards, and escalation paths • Create executive decision packages and run the cross-functional reliability council **Required Qualifications** • BS in electrical, mechanical, computer, materials, reliability engineering, physics, or equivalent experience (MS preferred) • 8+ years in hardware quality/reliability, server/rack systems, or mission-critical infrastructure • 3+ years owning field-failure, RMA, or CAPA outcomes • Solid understanding of hardware and system architecture across board, tray, rack, firmware, telemetry, and manufacturing test • Expertise in reliability statistics: censored life data, Weibull/Poisson/binomial methods, MTBF/MTTR, and reliability growth • Hands-on experience with FMEA/FTA, 8D/CAPA, FA, and corrective-action verification • Working proficiency with SQL and Python/R or equivalent analytics tools • Ability to influence design, validation, operations, suppliers/CMs, and senior leaders without direct authority **Preferred Skills** • GPU/AI server platforms, liquid cooling, high-power delivery, high-speed networking, or data-center operations • Design for serviceability and FRU strategy • Qualification-to-field correlation and mission-profile development • ODM/CM/supplier experience • Linux/BMC/IPMI/Redfish logs and fleet telemetry • Leadership of cross-generation reliability programs
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.