Staff Applied AI Inference Engineer
Crusoe · Denver, CO - US
About this role
## Staff Applied AI Inference Engineer Crusoe is on a mission to accelerate the abundance of energy and intelligence. As a vertically integrated AI infrastructure company, we own and operate each layer of the stack—from electrons to tokens—to power the world’s most ambitious AI workloads. We’re looking for problem-solving, opportunity-finding teammates with a sense of urgency who thrive in a fast-moving environment and want to grow their careers alongside experts across energy, manufacturing, data center construction, and cloud services. --- ## About the Role You’ll help make large language models run **faster, cheaper, and more reliably** in production by owning the inference stack end to end: - **Profile** where time and cost go - Bring **modern optimization techniques** into real deployments - Go deep into the **serving code** when defaults aren’t sufficient This is **core systems and performance work** on some of the most demanding models in use today. The work is applied—your optimizations land in **real customer deployments** with different models, traffic patterns, latency targets, and cost constraints. You’ll also work directly with customer engineering teams to tailor deployments to their needs, move workloads from proof of concept to fully monitored production services, and ensure the gains you engineer show up for the people running the workload. This is a **hands-on engineering role** centered on coding, profiling, and low-level optimization, with a customer-facing component and elements of product/technical solutions work. --- ## What You’ll Be Working On - Bring current inference techniques into production and refine them - Design and optimize serving architectures (e.g., **prefill/decode disaggregation**, request routing, and related approaches) - Work down into the serving stack—from **vLLM** and **SGLang** to **CUDA kernels**—profiling and analyzing to find and fix performance problems - Adapt and scale optimization methods across many ML model types, with an emphasis on **large language models** - Profile and tune deployments against targets for **latency, throughput, and cost**, and keep them dependable under real traffic - Tailor deployments to each customer’s models and constraints, partnering to move from proof of concept to live, well-monitored production - Build and support software/product features around the inference stack in production (Python preferred) - Experiment quickly
Listing freshness
CronJobs last confirmed this listing 17h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.