CronJobs

devops-sre jobs

Staff Technical Program Manager – Infrastructure Maintenance & Change Management

CoreWeave · Bellevue, WA

unknownsenior$177,000–$237,000Posted Oct 6, 2026PrometheusGrafanaOpenTelemetry

Apply on the employer site

About this role

**Staff Technical Program Manager – Infrastructure Maintenance & Change Management** **About the Role** CoreWeave is building the world’s largest AI Cloud platform, and our fleet is growing at extraordinary speed. We’re seeking a **Senior Technical Program Manager** to own **fleet-wide reliability and operations programs** that keep pace with that growth. You’ll partner closely with **Compute, Networking, Data Center, and Operations** teams to improve fleet operating efficiency—maximizing the percentage of fleet capacity that is **healthy, available, and sellable**—while driving **reliability and stability** goals. **What You’ll Do** - Establish and own **fleet reliability metrics and dashboards** (e.g., failure rates, MTTR, incident trends, capacity availability/utilization) and track how they evolve as the fleet expands. - Drive alignment on **fleet reliability OKRs** across engineering and operations. - Identify systemic reliability gaps across **hardware, firmware, software, networking, storage, and operational processes**. - Lead complex, cross-functional programs to improve **fleet delivery, readiness, and operational stability** at scale (solutions must hold as the fleet grows). - Define program plans, milestones, dependencies, risks, and success criteria for reliability initiatives. - Proactively manage cross-team dependencies and unblock execution across multiple engineering organizations. - Track progress against goals, surface risks early, and communicate status clearly to stakeholders and leadership. - Participate in major incident reviews and **root cause analysis**, ensuring follow-up actions are tracked to closure. - Use data and post-incident learnings to prioritize reliability investments and drive corrective action. **Minimum Qualifications** - Bachelor’s degree in Computer Engineering, Computer Science, or related field (or equivalent practical experience). - **7+ years** of technical program management experience in large-scale compute infrastructure, cloud, or platform environments. - Experience operating in rapidly growing/fast-scaling infrastructure environments, with programs designed to hold up under **2x, 5x, or greater** fleet growth. - Strong technical aptitude across infrastructure domains (compute, storage, networking, hardware, or SRE). - Demonstrated ability to use data and metrics to drive prioritization, execution, and decision-making. - Excellent communication and stakeholder management skills, including **executive-level reporting**. **Preferred Qualifications** - Experience at scale in data center, cloud infrastructure, or hyperscale environments—balancing reliability investments against capacity goals and understanding impact on sellable capacity. - Familiarity with reliability frameworks such as **SLIs/SLOs**, error budgets, incident management, and root cause analysis. - Understanding of hardware failure processes, including failure diagnosis, vendor coordination, and replacement lifecycle management. - Background in observability/monitoring/telemetry (e.g., **Prometheus, Grafana, OpenTelemetry**). - Experience with fleet lifecycle management (provisioning, firmware/OS updates, decommissioning) at scale. **Compensation & Benefits** - Base salary range: **$177,000 to $237,000** (final offer depends on knowledge, skills, experience, and market location). - Total rewards may include a discretionary bonus, equity awards, and benefits (eligibility-based). - US benefits (full-time) include: medica

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord