Staff Technical Program Manager – Infrastructure Maintenance & Change Management
CoreWeave · Bellevue, WA
About this role
**Staff Technical Program Manager – Infrastructure Maintenance & Change Management** **About the Role** CoreWeave is building the world’s largest AI Cloud platform, and our fleet is growing at extraordinary speed. We’re seeking a **Senior Technical Program Manager** to own **fleet-wide reliability and operations programs** that keep pace with that growth. You’ll partner closely with **Compute, Networking, Data Center, and Operations** teams to improve fleet operating efficiency—maximizing the percentage of fleet capacity that is **healthy, available, and sellable**—while driving **reliability and stability** goals. **What You’ll Do** - Establish and own **fleet reliability metrics and dashboards** (e.g., failure rates, MTTR, incident trends, capacity availability/utilization) and track how they evolve as the fleet expands. - Drive alignment on **fleet reliability OKRs** across engineering and operations. - Identify systemic reliability gaps across **hardware, firmware, software, networking, storage, and operational processes**. - Lead complex, cross-functional programs to improve **fleet delivery, readiness, and operational stability** at scale (solutions must hold as the fleet grows). - Define program plans, milestones, dependencies, risks, and success criteria for reliability initiatives. - Proactively manage cross-team dependencies and unblock execution across multiple engineering organizations. - Track progress against goals, surface risks early, and communicate status clearly to stakeholders and leadership. - Participate in major incident reviews and **root cause analysis**, ensuring follow-up actions are tracked to closure. - Use data and post-incident learnings to prioritize reliability investments and drive corrective action. **Minimum Qualifications** - Bachelor’s degree in Computer Engineering, Computer Science, or related field (or equivalent practical experience). - **7+ years** of technical program management experience in large-scale compute infrastructure, cloud, or platform environments. - Experience operating in rapidly growing/fast-scaling infrastructure environments, with programs designed to hold up under **2x, 5x, or greater** fleet growth. - Strong technical aptitude across infrastructure domains (compute, storage, networking, hardware, or SRE). - Demonstrated ability to use data and metrics to drive prioritization, execution, and decision-making. - Excellent communication and stakeholder management skills, including **executive-level reporting**. **Preferred Qualifications** - Experience at scale in data center, cloud infrastructure, or hyperscale environments—balancing reliability investments against capacity goals and understanding impact on sellable capacity. - Familiarity with reliability frameworks such as **SLIs/SLOs**, error budgets, incident management, and root cause analysis. - Understanding of hardware failure processes, including failure diagnosis, vendor coordination, and replacement lifecycle management. - Background in observability/monitoring/telemetry (e.g., **Prometheus, Grafana, OpenTelemetry**). - Experience with fleet lifecycle management (provisioning, firmware/OS updates, decommissioning) at scale. **Compensation & Benefits** - Base salary range: **$177,000 to $237,000** (final offer depends on knowledge, skills, experience, and market location). - Total rewards may include a discretionary bonus, equity awards, and benefits (eligibility-based). - US benefits (full-time) include: medica
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.