CronJobs

devops-sre jobs

Forward Deployed Engineer (Training)

Baseten · San Francisco

hybridunknown$200,000–$400,000Posted Aug 19, 2026KubernetesPyTorchJAXvLLMTensorRT-LLMSGLangRaySlurm

Apply on the employer site

About this role

## About Baseten Baseten powers mission-critical inference for the world’s most dynamic AI companies (like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer). By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we help frontier AI companies bring cutting-edge models into production. We’re growing quickly and recently raised our **$1.5B Series F**. Join us and help build the platform engineers turn to ship AI products. --- ## The Role — Forward Deployed Engineer (Training) Forward Deployed Engineers work directly with the largest and fastest-growing AI companies in the world, owning technical outcomes on Baseten and tackling the hardest problems in serving and improving models at scale. You’ll work across the model lifecycle: **inference**, **post-training**, and the systems that tighten the loop between them. ### What you’ll do - Act as each account’s **de facto CTO** on Baseten, with final accountability for how workloads are designed, run, and scaled. - Take customer objectives from vague to shipped: frame the problem, define specs and success criteria, build a PoC, and move it to production quickly. - Design **evals and benchmarks** to isolate quality/performance gaps, then close the gap (e.g., optimize inference, improve models via post-training, or rework the eval). - Be the first responder to mission-critical failures: triage, fix directly or route to the owning team, and stay accountable until it ships. - Build internal systems to make each engagement faster: tooling/automation for eval and deployment infrastructure, plus recipes and reference implementations. - Shape the product by channeling account needs into the roadmap and shipping fixes/features into Baseten’s codebase. - Manage multiple accounts at once—sequencing work, pulling in the right people at the right time, and keeping stakeholders aligned on status and risk. --- ## Requirements - **1–2+ years** of software engineering experience shipping and maintaining code in large production systems (ideally across the stack) - Experience debugging complex production issues (logs, metrics, traces, root-cause in unfamiliar systems) - Confidence owning ambiguous technical problems (triage, decisions under uncertainty, knowing when to pull in others) - Motivation beyond pure engineering: interest in working directly with customers and influencing the product - Clear communication on complex technical topics (with both engineers and leadership) - Genuine curiosity about AI inference/training and a drive to become an expert in the infrastructure powering it - Willingness to respond to customers outside regular hours and participate in an on-call rotation - Excitement about solving problems for major AI companies running mission-critical workloads --- ## What You’ll Bring (Optional Strengths) We don’t expect one person to cover everything—strong candidates may have depth in one or more areas: - Core infrastructure domains (e.g., storage systems, networking—from cloud to cluster interconnects like InfiniBand/RoCE) - Operating distributed compute platforms (e.g., Kubernetes, Slurm, Ray), especially for GPU workloads - Understanding of LLM architectures and inference engines (e.g., vLLM, TensorRT-LLM, SGLang) - Profiling and optimizing GPU workloads (training or serving) - Post-training experience (e.g., SFT, RL) and/or deep learning background with PyTorch or JAX - Operational depth (on-call, incident response, debugging distr

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord