Senior Software Engineer, Model Infrastructure
Harvey · San Francisco
About this role
**Senior Software Engineer, Model Infrastructure — Harvey** **Why Harvey** At Harvey, we’re transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. We’re scaling fast, defining a new category in real time, and building a generational company at a true inflection point. We move quickly, take ownership, and stay close to customers—guided by three values: **Decisiveness, Simplicity, and Job’s Not Finished**. --- **Role Overview** As a **Staff Software Engineer** on the **Model Infrastructure** team, you’ll lead the design and development of the systems that power **every AI request** at Harvey. You’ll partner with **AI Research, Product Engineering, Infrastructure**, and **external model providers** to build a platform that is **reliable, scalable, observable, and efficient**. --- **What You’ll Build** - **Model Reliability & Operations** - Model health monitoring - Automated failover and recovery - Capacity provisioning - Operational tooling and incident automation - **Unified Model Controller (UMC)** - Policy-based model routing - Intelligent model selection - Traffic management - Reliability and latency optimization - **Provider Platform** - Multi-provider architecture - API and SDK integrations (e.g., OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, and future providers) - Rapid adoption of new frontier models - **Observability & Cost Platform** - Token usage analytics - Cost attribution - Latency and reliability dashboards - Capacity forecasting - Utilization optimization - **AI Platform Foundation** - Infrastructure supporting model evaluation - Model deployment and operations - Future model training platform - Agent infrastructure and CcaaS --- **What You’ll Do** - Lead the design and implementation of Harvey’s **Model Infrastructure** platform - Build systems that ensure **high availability**, **low latency**, and **operational excellence** for AI inference - Design and improve the **Unified Model Controller (UMC)** and **Model Selector** to detect degradations and route traffic based on **reliability, latency, quality, compliance, and cost** - Develop systems for **model provisioning, capacity management, failover, and traffic engineering** across multiple AI providers - Integrate new model providers and maintain provider APIs/SDKs to enable rapid adoption of emerging frontier models - Improve observability with health dashboards, alerting, token usage analytics, cost reporting, and end-to-end telemetry - Partner with Product Engineering to support model launches, experimentation, and proactive monitoring of production AI workloads - Drive infrastructure efficiency through capacity planning, utilization optimization, and cost visibility --- *Note: The provided description appears truncated at the end.*
Listing freshness
CronJobs last confirmed this listing 15h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.