Software Engineer - Voice AI (Inference Runtime)
Baseten · San Francisco
About this role
**Software Engineer - Voice AI (Inference Runtime)** **About Baseten** Baseten powers mission-critical inference for leading AI companies like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. Recently raised $1.5B Series F led by Altimeter Capital, Conviction Partners, and Spark Capital. **The Role** Join a small founding team building production-grade Voice AI infrastructure. You'll own Baseten Voice AI's inference stack end-to-end—from product roadmap through engineering implementation. Partner with Forward Deployed Engineers and Model Performance Engineers to push Voice AI boundaries across productivity, customer service, clinical conversation, creator tools, and education. **Key Initiatives** • Develop world-class model serving stack for open-source voice models—reduce latency, increase throughput, improve GPU efficiency • Build large-scale real-time infrastructure for multi-model voice agents with streaming I/O • Design training/inference iteration loops for voice model customization • Past wins: World's fastest Whisper with streaming/diarization; Canopy Labs TTS inference partnership **Responsibilities** • Own Voice AI product areas end-to-end: architecture, design, implementation, rollout, operations • Design and operate real-time, large-scale systems for STT, TTS, and voice agent workloads • Drive cross-team collaboration on full-stack technical problems • Mentor teammates through code reviews and technical leadership **Requirements** • Bachelor's in Computer Science or related field • Proven track record owning production real-time systems where p99 latency matters • Proficient in Python or similar languages • Strong product taste, especially for developer tools • Interest in ML/AI infrastructure • Comfortable using AI coding assistants (Claude, Cursor, Codex) daily • Strong collaboration and communication skills **Nice to Have** • Pipeline-level runtime optimizations (dynamic batching, async scheduling) • Developer platform experience (SDKs, CLIs, APIs) • Docker, Kubernetes, service meshes, distributed scheduling • Speech/audio ML models (STT, TTS) • Model-serving runtimes (vLLM, TensorRT, ONNX) • Systems-level performance profiling and GPU diagnostics • Customer-facing engineering experience **Benefits** • Competitive compensation with meaningful equity • 100% medical, dental, vision coverage (U.S.) • Flexible PTO + company Winter Break • Paid parental leave • Fertility/family-building stipend (Carrot) • 401(k) facilitation (U.S.) • Exposure to ML startup ecosystem **Equal Opportunity Employer** | Committed to diversity and inclusion
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.