CronJobs

devops-sre jobs

Senior Manager, Production & Fleet Operations

Cerebras · Sunnyvale, CA

remoteseniorPosted Oct 9, 2026

Apply on the employer site

About this role

## About Cerebras Cerebras Systems builds the world’s largest AI chip—an architecture designed to deliver industry-leading training and inference speeds (reported as **10x+ faster than GPU-based hyperscale cloud inference**). Cerebras works with leading model labs, global enterprises, and AI-native startups, including a recent multi-year partnership with OpenAI to deploy **750MW** of scale. --- ## The Role Cerebras is building and operating advanced AI infrastructure. As the infrastructure footprint expands across **cloud, colocation facilities, and customer environments**, the team needs a scalable operating model with continuous visibility into production infrastructure—so issues can be **identified, prioritized, and resolved quickly**. You will lead the **day-to-day operational management** of Cerebras-managed production infrastructure, reporting to the **Director of Central Operations**. You’ll establish and operate mechanisms to understand **fleet health**, coordinate **production response**, manage **maintenance and repair**, and ensure infrastructure is **safely and efficiently returned to service** after failures. You’ll work closely with: **SiteOps/DC Ops, Reliability & Incident Management, Service Ops & Enablement, Global Service Logistics & Inventory, Build & Deploy, and Engineering**. --- ## Mission Operate the Cerebras production fleet **safely, reliably, and consistently**—providing: - Continuous visibility into infrastructure health - Rapid restoration when failures occur - An increasingly standardized and automated operating model as the fleet scales --- ## Responsibilities ### 1) Own Production Fleet Operations - Own the day-to-day operational health of Cerebras-managed production infrastructure - Establish consistent operational processes across clusters, sites, and operating environments - Maintain clear visibility into production state, degradation, outages, maintenance, and infrastructure availability - Establish fleet monitoring, alert response, triage, escalation, and restoration processes - Develop appropriate **24x7 operational coverage** as fleet requirements evolve - Establish clear ownership of production issues from detection through restoration ### 2) Build the Fleet Operations Capability - Develop and mature the **Fleet Operations / NOC** capability - Establish standards for monitoring, telemetry, alerting, dashboards, and fleet-health reporting - Develop shift, handoff, escalation, and on-call operating procedures - Ensure operators have the tooling, runbooks, procedures, and access required to safely operate production infrastructure - Partner with Service Ops & Enablement to develop operator competency and training requirements - Drive standardization across geographically distributed infrastructure ### 3) Coordinate Physical Intervention - Determine when physical intervention is required - Dispatch and prioritize SiteOps/DC Ops activities based on production impact - Provide clear diagnostic information and execution procedures - Coordinate troubleshooting between centralized Operations and (additional teams as needed)

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord