CronJobs

devops-sre jobs

Senior Hardware Engineer, Server Infrastructure

CoreWeave · New York, NY/ Bellevue, WA

remoteseniorPosted Sep 29, 2026PythonAnsibleRedfishIPMIKubernetesPrometheusGrafanaLinux

Apply on the employer site

About this role

**Senior Hardware Engineer, Server Infrastructure** **About the Role** CoreWeave is seeking a highly skilled and motivated engineer to join our Hardware Engineering team. In this role, you will design, develop, and optimize CoreWeave’s server hardware infrastructure—owning hardware from provisioning through decommission. This is a hands-on position combining engineering and operational support, including automation across the hardware lifecycle, hardware/firmware management, monitoring and alerting, qualification and bring-up of new platforms, and deep root-cause analysis for hardware escalations. **What You’ll Do** - Design and develop server hardware infrastructure to support high-performance workloads - Automate the server hardware lifecycle (provisioning/configuration → firmware management → monitoring → decommissioning) - Develop and maintain hardware and firmware management services for reliability at scale - Build monitoring and alerting for server hardware health; improve alert quality for reliable on-call response - Serve as a senior point of contact for hardware escalations; perform deep troubleshooting and root-cause analysis across hardware and firmware - Participate in an on-call rotation and improve runbooks, alerts, and tooling - Collaborate with cross-functional teams to define hardware requirements, specifications, and system architecture - Work with server vendors and OEMs to evaluate, qualify, deploy new platforms, and resolve firmware/quality/RMA issues - Support new data center region bring-up and hardware qualification - Analyze performance, identify bottlenecks, and implement improvements to efficiency and resilience - Establish and refine processes for internal hardware testing, deployment, and performance optimization - Create and maintain documentation (designs, specs, test procedures, results) - Support data center operations teams and hardware technicians with troubleshooting guidance, runbooks, and training - Turn recurring production failures into automation, better telemetry, and platform/vendor improvements - Communicate status, trade-offs, and risks clearly during incidents and across stakeholders **Who You Are** - Deep understanding of server hardware, components, and management technologies - Proficiency in Ansible or Python; hands-on experience programmatically interacting with server BMCs using Redfish or IPMI (Redfish preferred) - Experience collaborating with hardware vendors/OEMs to evaluate, qualify, and deploy server solutions - Proven experience supporting and troubleshooting production infrastructure; on-call/escalation experience - Comfortable building/automating systems and supporting infrastructure already in production - Ability to stay current with server and data center hardware technologies and trends - Strong interest in automation and infrastructure scalability; commitment to continuous improvement - Excellent technical documentation skills and attention to detail - Strong analytical/problem-solving skills with a systematic, data-driven approach - Excellent written and verbal communication skills in English **Preferred Qualifications** - Experience bringing up new data center regions / standing up infrastructure in new geographies - Experience with GPU platforms and rack-scale systems (e.g., NVIDIA GB200/GB300) - Experience with BMC technologies at fleet scale (Redfish/IPMI) - Experience with firmware lifecycle management and/or hardware qualification programs in large-scale envir

Listing freshness

CronJobs last confirmed this listing 52m ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord