Senior Manager, Engineering - AI Inference
Crusoe · San Francisco, CA - US
About this role
**Senior Manager, Engineering — AI Inference (Crusoe)** Crusoe is on a mission to accelerate the abundance of energy and intelligence. As a vertically integrated AI infrastructure company, we build and operate the full stack—from electrons to tokens—to power the world’s most ambitious AI workloads. We’re looking for an engineering leader who can drive performance and reliability improvements for large language model inference in production—while also partnering with customers to tailor deployments to real-world constraints. --- ## About the Role As a **Senior Engineering Manager**, you will lead an engineering team focused on making large language models **run faster, cheaper, and more reliably** in production. You’ll balance **people leadership** with **hands-on technical work** across the inference stack end to end—profiling performance, applying modern optimization techniques, and diving into serving code (including kernel-level performance analysis). This is **applied engineering**: optimizations you build should land in **real customer deployments**, each with unique models, traffic patterns, latency targets, and cost constraints. --- ## What You’ll Be Working On - Bring current **inference techniques** into production and refine them - Design and optimize serving architectures (e.g., **prefill/decode disaggregation**, request routing) - Work deep in the serving stack—from **vLLM/SGLang** down to **CUDA kernels**—to profile and fix performance issues - Adapt and scale optimization methods across many ML model types, with emphasis on **LLMs** - Profile and tune deployments against targets for **latency, throughput, and cost**, ensuring dependability under real traffic - Tailor deployments to each customer’s models and constraints; move workloads from **proof of concept → monitored production** - Build and support production software/product features around the inference stack (Python preferred) - Run fast experiments: turn fuzzy goals into specs, execute focused proofs of concept, and ship well-tested results - Lead delivery end to end: set performance goals, guide projects to production, and partner on technical strategy/roadmaps - Make sound tradeoffs in ambiguous environments; avoid unnecessary complexity - Take ownership and accountability for outcomes and team execution --- ## What You’ll Bring to the Team - **2+ years** managing and leading an engineering team in a high-performance or ML-focused environment - Strong hands-on experience in software engineering, low-level optimization, or ML infrastructure—and a desire to stay close to code/architecture - BS/MS/PhD in CS, Engineering, Math, or related field - Production shipping experience in general-purpose languages (Python preferred; C++ also relevant) - Familiarity with methods for **high-throughput / low-latency LLM inference** - Comfort with modern LLM serving frameworks (**vLLM** or **SGLang**) and kernel-level performance profiling - Strong understanding of **GPU behavior** - Interest and hands-on experience with **large language models** - Working knowledge of AI/ML pipelines from development through deployment - Strong communication skills, including explaining complex technical topics to customers and teammates ### Bonus Points - Track record of making systems run faster (especially for LLMs) - Experience with **CUDA** (or comparable technologies) - Strong AI/ML inference system fundamentals and shipping experience - Experience with **Docker** and **Kuberne
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.