Capacity Operations Manager
Baseten · San Francisco
About this role
## About Baseten Baseten powers mission-critical inference for the world’s most dynamic AI companies (e.g., Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer). By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable frontier AI companies to bring cutting-edge models into production. We’re growing quickly and recently raised our **$1.5B Series F** (led by Altimeter Capital, Conviction Partners, and Spark Capital). Join us and help build the platform engineers turn to ship AI products. ## The Role We’re looking for a **hands-on Operations Manager** to own the **operational and analytical supply side** of our **GPU fleet**. **Key focus areas:** GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our **neocloud** and **bare metal** environments. This is an **operator role** (not people management). You’ll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment. ## Responsibilities - Drive suppliers to keep the **maximum amount of the GPU fleet** online and healthy. - Maintain a live reconciliation of **contracted vs. provisioned vs. healthy vs. utilized** capacity, broken out by **supplier** and by **cluster**, maximizing the number of healthy GPUs. - Supplier-attributed fleet health accountability: own **replacement SLAs**, **mean time to repair (MTTR)**, and **RMA cycle times** for every in-scope supplier. - **SLA monitoring, credit claims, and remedy enforcement:** track SLA performance vs. contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short. - Drive internal communications when suppliers need to perform maintenance so all Baseten stakeholders are aware of activities impacting availability. ## Scope & Approach - **Flexibility:** this list covers the core of the role, not the limit of it. - **Ownership mindset:** “whatever it takes” is a real operating principle—if it’s outside a defined lane but inside the goal of closing the capacity gap, it’s yours to pick up. ## Requirements - **5 to 10+ years** in infrastructure, working within the **compute lifecycle** to maximize functional compute (ideally in a hyperscale, cloud, or large-scale compute environment). - Direct experience managing **GPU/server/data center hardware supplier relationships**. - You understand fleet health, **RMA processes**, and how **contracted capacity** differs from **delivered capacity**. - Highly analytical: comfortable pulling your own data, building your own reports, and generating insights without waiting for someone else’s dashboard. - Comfortable with ambiguity: part of the job is figuring out what should exist and building it. - Strong cross-functional collaboration: you’ll work closely with **finance, infrastructure/engineering, legal, and security**. ## Nice to Have - Experience at a hyperscaler, neo cloud provider, or AI infrastructure company. - Familiarity with GPU hardware lifecycles (e.g., **NVIDIA H100/H200/GB200 class systems**), power/thermal constraints, and compute supply chain dynamics. - Experience running formal supplier corrective actions. ## Benefits - Competitive compensation, including meaningful equity. - **(U.S. only)** 100% coverage of medical, dental, and vision insurance for employee and dependents. - Flexible PTO policy, including company-wide **Winter Break** (offices closed from Christmas Eve to New Year’s Day). - Paid paren
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.