CronJobs

devops-sre jobs

Data Center Hardware Quality & Reliability Engineer

OpenAI · San Francisco

hybridsenior$226,000–$285,000Posted Sep 28, 2026PythonSQLLinuxIPMIRedfish

Apply on the employer site

About this role

**Data Center Hardware Quality & Reliability Engineer** **About The Role** Own the end-to-end data-center hardware quality and reliability loop for OpenAI's 3P infrastructure and 1P current and next-gen platforms. Turn field failures into quantified risk, fast containment, verified root cause, improved MQE/NPI and manufacturing-test coverage, accurate spares forecasts, and upstream changes that prevent recurrence. **Key Responsibilities** • Build and govern the field-quality data model across telemetry, tickets, RMA/repair, FA, firmware, configuration, supplier, and manufacturing genealogy • Define AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF, repair-cycle-time, and forecast-versus-actual metrics • Provide fleet-level macro views and unit/FRU/cohort-level micro views; detect shifts and bound affected populations • Lead systemic field-failure triage, containment, failure analysis, 8D/CAPA, risk assessment, and corrective-action verification • Develop cohort, life-data, reliability-growth, and spare-demand projections by product, FRU, supplier, configuration, geography, and age • Partner with MQE and NPI to convert field mechanisms into manufacturing-test coverage and qualification gates • Define supplier/CM FA standards, field-data contracts, scorecards, and escalation paths • Create executive decision packages and run the cross-functional reliability council **Required Qualifications** • BS in electrical, mechanical, computer, materials, reliability engineering, physics, or equivalent experience (MS preferred) • 8+ years in hardware quality/reliability, server/rack systems, or mission-critical infrastructure • 3+ years owning field-failure, RMA, or CAPA outcomes • Solid understanding of hardware and system architecture across board, tray, rack, firmware, telemetry, and manufacturing test • Expertise in reliability statistics: censored life data, Weibull/Poisson/binomial methods, MTBF/MTTR, and reliability growth • Hands-on experience with FMEA/FTA, 8D/CAPA, FA, and corrective-action verification • Working proficiency with SQL and Python/R or equivalent analytics tools • Ability to influence design, validation, operations, suppliers/CMs, and senior leaders without direct authority **Preferred Skills** • GPU/AI server platforms, liquid cooling, high-power delivery, high-speed networking, or data-center operations • Design for serviceability and FRU strategy • Qualification-to-field correlation and mission-profile development • ODM/CM/supplier experience • Linux/BMC/IPMI/Redfish logs and fleet telemetry • Leadership of cross-generation reliability programs

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord