Software Engineer - AI Developer Productivity
Baseten · San Francisco
About this role
**ABOUT BASETEN** Baseten powers mission-critical inference for the world’s most dynamic AI companies (e.g., Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, Writer). By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we help frontier AI companies bring cutting-edge models into production. We’re growing quickly and recently raised our **$1.5B Series F**. Join us to help build the platform engineers turn to when shipping AI products. --- **THE ROLE** Baseten engineers want to work in an **AI-first** way. What’s missing isn’t enthusiasm—it’s the **platform underneath it**. Today, teams assemble their own agent configs, context files, and MCP servers. Your job is to build the shared platform so the best patterns become **defaults everyone inherits**. You’ll build: - **Agent configurations** tuned to our monorepo - A **context + tooling layer** that makes agents competent in our codebase - **Evals** to measure what actually works - **Rollout mechanics** that get new engineers productive with agents in week one Success looks like teams adopting what you build because it’s better than what they’d cobble together—not because a policy forces it. --- **EXAMPLE INITIATIVES** - **Agent substrate**: repo-level context infrastructure (e.g., CLAUDE.md/AGENTS.md conventions, architecture/domain context, tooling to keep it accurate as code changes) - **Internal MCP servers**: scoped access to CI, observability, incident tooling, deployment state, and docs - **Shared skills/subagents/hooks**: encode Baseten workflows - **Sandboxed environments**: safe build + test for agents - **Golden path**: project templates and onboarding with AI tooling configured and working - **Self-serve infrastructure**: teams can build their own agents without you as a bottleneck - **Gateway/auth/cost controls/audit logging** for internal model access - **Feedback loop**: eval harnesses comparing configs on real Baseten tasks (not vibes) - **Instrumentation**: measure AI tool usage and downstream effects (cycle time, review latency, change failure rate) - **Agents in the SDLC**: PR review triage, test gap-filling, incident context assembly, migrations/refactors, codebase Q&A - **CI/CD integration** with guardrails to make agent workflows trustworthy --- **RESPONSIBILITIES** - Own the internal AI developer platform end-to-end: architecture, build, rollout, operation, measurement - Evaluate and integrate third-party AI coding tools (e.g., Claude Code, Cursor, Codex) and build the context layer for our monorepo - Build frameworks so engineers can create their own agents without deep LLM expertise - Establish and drive AI-assisted development evaluation practices - Drive adoption via developer experience (good defaults, clear docs, low friction—not mandates) - Embed with teams to identify where AI unblocks work, then generalize into platform capabilities - Own the safety layer: permissions, secrets handling, audit trails, cost management --- **REQUIREMENTS** - **4+ years** building and enabling AI-native SDLC - Strong proficiency in **Python and/or Go**; build tools other engineers rely on daily - Hands-on experience with **LLMs + agent frameworks** (tool calling, MCP, context management, orchestration, failure handling) - Shipped agentic work used by real people (not just prototypes) - Deep fluency with AI coding tools and strong opinions about where they break down - **Platform mindset**: build for adoption and self-service; t
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.