CronJobs

devops-sre jobs

Technical Program Manager, Model Performance

Baseten · San Francisco

hybridunknown$165,000–$330,000Posted Aug 29, 2026vLLMTensorRT-LLMSGLangNVIDIA Dynamodeep learninginference optimization

Apply on the employer site

About this role

**ABOUT BASETEN** Baseten powers mission-critical inference for the world’s most dynamic AI companies (e.g., Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, Writer). By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we help frontier AI companies bring cutting-edge models into production. Baseten is growing quickly and recently raised a **$1.5B Series F** (led by Altimeter Capital, Conviction Partners, and Spark Capital). Join us to help build the platform engineers turn to ship AI products. --- **THE ROLE** The **Model Performance** organization at Baseten is hiring its **first Technical Program Manager**. This is a **zero-to-one** role in a team responsible for building the core algorithms and methods that power Baseten’s high-performance inference stack. You won’t inherit an existing program framework—you’ll build one from the ground up: **planning structure, execution processes, metrics, and cross-functional alignment**. If you can take ambitious but loosely defined initiatives and turn them into a **predictable, well-governed program**, this role is for you. **Example initiatives (from the team):** - How to build a day-0 API for Kimi K3: https://www.baseten.co/blog/how-to-build-a-day-zero-api-for-kimi-k3/ - How we built the new fastest API for GLM-5.2: https://www.baseten.co/blog/how-we-built-the-new-fastest-api-for-glm-52/ - Inference engineering for DeepSeek V4 Pro 0813: https://www.baseten.co/blog/inference-engineering-for-deepseek-v4-pro-0813/ --- **RESPONSIBILITIES** - Own execution across Model Performance’s active project portfolio, freeing technical leads to focus on technical direction. - Design and stand up planning structures, operating cadences, and status reporting mechanisms. - Coordinate model release and optimization programs end to end, including day-zero launches and sequencing across performance engineering, infra, and release stakeholders. - Drive cross-team alignment as scope expands from Model Performance Core into Model APIs and the inference production stack (BIS). - Surface risks and dependencies early; keep leadership informed with clear, honest status. - Partner with engineering leads on team structures and ownership boundaries as the org scales. --- **REQUIREMENTS** - Deep technical program management experience; you’re already running programs of similar or greater complexity. - Experience program-managing model performance or inference optimization work. - You understand how engines like **vLLM, TensorRT-LLM, SGLang, or NVIDIA Dynamo** fit into a production serving stack and can engage credibly with engineers. - Comfort with ambiguity and **zero-to-one** program building. - Proven ability to influence without authority across engineers, managers, and leadership. - Excellent written and verbal communication; able to make deeply technical programs legible to any audience. - High agency decision-making with strong ownership and accountability. --- **NICE TO HAVE** - Experience coordinating model release programs, including day-zero launches (ideally for large-scale models). - Prior hands-on software engineering background. - Deep learning performance optimization background (training or inference; inference preferred). --- **BENEFITS** - Competitive compensation, including meaningful equity - **(U.S. only)** 100% coverage of medical, dental, and vision insurance for employee and dependents - Flexible PTO policy, including company-wide Winter Break

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord