CronJobs

backend jobs

Inference Engineer

Cartesia · *HQ - San Francisco, CA

onsiteunknown$180,000–$250,000Posted Dec 12, 2024PythonCUDATritonvLLMSGLangTransformersDistributed SystemsMachine Learning

Apply on the employer site

About this role

**Inference Engineer** **About Cartesia** Pioneering AI that learns from and interacts with the world like humans do. Founded by Stanford AI Lab PhDs who invented State Space Models (SSMs). Funded by Index Ventures, Lightspeed Venture Partners, and 90+ industry experts. **Your Impact** • Design and build low-latency, scalable inference and serving stack for foundation models • Work with research and product teams to serve cutting-edge AI products efficiently • Build robust inference infrastructure and monitoring systems • Shape products with significant autonomy and direct impact **What You Bring** • Strong engineering skills with clean, maintainable code • Experience building large-scale distributed systems with high performance demands • Technical leadership and ability to execute zero-to-one results • Background in inference pipelines with ML and generative models • Experience implementing state-of-the-art ML models in production • *Preferred:* vLLM, SGLang, Continuous Batching, or similar frameworks • *Preferred:* CUDA, Triton, or similar experience **Details** 📍 **Locations:** San Francisco, London, Bangalore (in-person) 🌍 **Visa Sponsorship:** Case-by-case basis ⚡ **Culture:** Fast-shipping, high-bar execution, supportive team **Benefits (US Employees)** 💰 Competitive salary + equity | 🩺 Full health coverage | 👨‍👩‍👧‍👦 9-12 weeks parental leave | 🏦 401(k) | 🚆 Commuter allowance | 🏖️ Flexible PTO | 🍲 Daily meals & snacks *Equal opportunity employer*

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord