CronJobs

devops-sre jobs

Sr./Staff TPM - Inference Capacity

Cerebras · Sunnyvale, CA

hybridseniorPosted Jun 29, 2026PythonSQLGrafanaFluxJiraConfluenceKV cacheaccelerator scheduling

Apply on the employer site

About this role

## About Cerebras Cerebras Systems builds the world’s largest AI chip—an architecture designed to deliver industry-leading training and inference speeds (over 10× faster than GPU-based hyperscale cloud inference services). Cerebras works with leading model labs, global enterprises, and AI-native startups, including a recent multi-year partnership with OpenAI to deploy 750MW of scale. ## About the Role As demand for AI accelerates, intelligent capacity management is one of the company’s most strategic challenges. You’ll lead capacity planning and fleet strategy for the **Inference Service** organization—working closely with Engineering, Product, Infrastructure, SRE, Operations, and executive leadership to maximize utilization of advanced AI inference fleets. ## What You’ll Own - **Capacity planning & forecasting**: Build and maintain a **6 / 12 / 26-week rolling capacity model** across every cluster; translate customer contracts and sales pipeline asks into capacity requirements; forecast replicas, system-hours, and spares by customer and model; reconcile against actuals weekly; maintain the source-of-truth documentation. - **New datacenter capacity bring-up**: Partner with datacenter infrastructure and operations teams to support new datacenter bring-up and ensure production readiness; drive engineering efforts and automation for on-time, high-quality delivery. - **Allocation & cluster placement**: Run weekly capacity reviews across customers/models/clusters; decide model placement and re-balancing (tenant-to-cluster mapping, which clusters absorb launches, freezes in effect, etc.); publish weekly capacity/utilization reporting for Inference Service leadership; coordinate downstream deployment tasks with the SRE team. - **Capacity planning tool adoption**: Partner with console engineering to drive stakeholder adoption of the in-house capacity planning/allocation tool, including UAT, issue resolution, change tracking, pilot testing, and deployment. - **Process improvement**: Contribute to continuous improvement of internal capacity management tools. - **Incident tracking & postmortems**: Proactively identify and mitigate capacity bottlenecks, risks, and dependencies; if SLAs drop due to misallocation, drive resolution and postmortem. ## Key Responsibilities - Run **weekly capacity planning** and **daily capacity/deployment tracking** with Engineering, Product, and Operations. - Own **fleet utilization reporting and forecasting**. - Drive capacity planning for **new customer deployments** and **major model launches**. - Lead continuous improvement and stakeholder adoption of the capacity management platform. - Drive org-level strategic initiatives for capacity expansion, fleet efficiency, and maximizing utilization. - Lead planning for major infrastructure events (e.g., new customer commits, new model releases, DC/cluster architecture changes) and update forecasts accordingly. - Maintain **Jira EPICs** and **Confluence pages** for capacity planning, reporting, and change management to ensure execution transparency. ## Qualifications - **5+ years** in TPM / technical program management / product operations experience in cloud infrastructure, large-scale ML serving, or hyperscaler capacity planning. - Experience leading large cross-functional programs across **Engineering, Product, and Operations**. - Comfort with the inference serving stack (e.g., model replicas, batching, prefill/decode, KV cache, accelerator scheduling). - Strong dat

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord