AI Inference Engineer
Baseten · San Francisco
About this role
**About Baseten** Baseten powers mission-critical inference for the world’s most dynamic AI companies (e.g., Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, Writer). By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we help frontier AI companies bring cutting-edge models into production. We’re growing quickly and recently raised our **$1.5B Series F** (led by Altimeter Capital, Conviction Partners, and Spark Capital). Join us and help build the platform engineers turn to ship AI products. --- **The Role — Forward Deployed Engineer** As a Forward Deployed Engineer at Baseten, you’ll partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten’s platform. You’ll own the journey from initial exploration to production deployment—translating ambiguous business goals into reliable, observable services with clear **quality, latency, and cost** outcomes. This is a hands-on engineering role with coding and software development, plus product-management and customer-facing responsibilities (technical customer success and pre-sales solution engineering). --- **Example Initiatives** - Forward Deployed Engineering on the frontier of AI: https://www.baseten.co/blog/forward-deployed-engineering/ - The fastest, most accurate Whisper transcription: https://www.baseten.co/blog/the-fastest-most-accurate-and-cost-efficient-whisper-transcription/ - Deploy production-ready model servers from Docker images: https://www.baseten.co/blog/deploy-production-model-servers-from-docker-images/ - Deploy custom ComfyUI workflows as APIs: https://www.baseten.co/blog/deploying-custom-comfyui-workflows-as-apis/ --- **Responsibilities** - Develop and maintain production software systems and product features (preference for **Python**). - Drive customer impact end-to-end: **problem framing → evaluation → production deployment → monitoring**. - Deliver with velocity: turn vague objectives into clear specs and well-defined PoCs to ship well-tested services. - Optimize and enhance AI/ML projects and contribute to continuous improvement of the technical stack (including features and PRDs). - Own products and customer projects end-to-end—acting as engineer, project manager, and product manager. - Navigate ambiguity and make good tradeoffs without unnecessary complexity. - Demonstrate ownership and accountability. --- **Requirements** - BS/MS/PhD in Computer Science, Engineering, Mathematics, or related field. - **2+ years** of professional experience in a fast-paced, high-growth environment. - Production experience with one or more general-purpose languages (strong preference for **Python**). - Familiarity with AI/ML pipelines and the ML model development + deployment lifecycle. - Strong communication skills on complex technical topics. - Experience building or optimizing AI/ML projects is highly valued. --- **Benefits** - Competitive compensation, including meaningful equity - **(U.S. only)** 100% coverage of medical, dental, and vision insurance for employee and dependents - Flexible PTO, including company-wide Winter Break (offices closed **Dec 24–Jan 1**) - Paid parental leave - Fertility and family-building stipend through Carrot - **(U.S. only)** Company-facilitated 401(k) - Exposure to a variety of ML startups for learning and networking --- **Equal Opportunity** Baseten is committed to fostering a diverse and inclusive workplace and is an Equal Opportunity Employe
Listing freshness
CronJobs last confirmed this listing 3h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.