Sr./Staff TPM - Inference Capacity
Cerebras · Sunnyvale, CA
About this role
## About Cerebras Cerebras Systems builds the world’s largest AI chip—an architecture designed to deliver industry-leading training and inference speeds (over 10× faster than GPU-based hyperscale cloud inference services). Cerebras works with leading model labs, global enterprises, and AI-native startups, including a recent multi-year partnership with OpenAI to deploy 750MW of scale. ## About the Role As demand for AI accelerates, intelligent capacity management is one of the company’s most strategic challenges. You’ll lead capacity planning and fleet strategy for the **Inference Service** organization—working closely with Engineering, Product, Infrastructure, SRE, Operations, and executive leadership to maximize utilization of advanced AI inference fleets. ## What You’ll Own - **Capacity planning & forecasting**: Build and maintain a **6 / 12 / 26-week rolling capacity model** across every cluster; translate customer contracts and sales pipeline asks into capacity requirements; forecast replicas, system-hours, and spares by customer and model; reconcile against actuals weekly; maintain the source-of-truth documentation. - **New datacenter capacity bring-up**: Partner with datacenter infrastructure and operations teams to support new datacenter bring-up and ensure production readiness; drive engineering efforts and automation for on-time, high-quality delivery. - **Allocation & cluster placement**: Run weekly capacity reviews across customers/models/clusters; decide model placement and re-balancing (tenant-to-cluster mapping, which clusters absorb launches, freezes in effect, etc.); publish weekly capacity/utilization reporting for Inference Service leadership; coordinate downstream deployment tasks with the SRE team. - **Capacity planning tool adoption**: Partner with console engineering to drive stakeholder adoption of the in-house capacity planning/allocation tool, including UAT, issue resolution, change tracking, pilot testing, and deployment. - **Process improvement**: Contribute to continuous improvement of internal capacity management tools. - **Incident tracking & postmortems**: Proactively identify and mitigate capacity bottlenecks, risks, and dependencies; if SLAs drop due to misallocation, drive resolution and postmortem. ## Key Responsibilities - Run **weekly capacity planning** and **daily capacity/deployment tracking** with Engineering, Product, and Operations. - Own **fleet utilization reporting and forecasting**. - Drive capacity planning for **new customer deployments** and **major model launches**. - Lead continuous improvement and stakeholder adoption of the capacity management platform. - Drive org-level strategic initiatives for capacity expansion, fleet efficiency, and maximizing utilization. - Lead planning for major infrastructure events (e.g., new customer commits, new model releases, DC/cluster architecture changes) and update forecasts accordingly. - Maintain **Jira EPICs** and **Confluence pages** for capacity planning, reporting, and change management to ensure execution transparency. ## Qualifications - **5+ years** in TPM / technical program management / product operations experience in cloud infrastructure, large-scale ML serving, or hyperscaler capacity planning. - Experience leading large cross-functional programs across **Engineering, Product, and Operations**. - Comfort with the inference serving stack (e.g., model replicas, batching, prefill/decode, KV cache, accelerator scheduling). - Strong dat
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.