Software Engineer - Training Infrastructure
Baseten · San Francisco
About this role
**ABOUT BASETEN** Baseten powers mission-critical inference for the world’s most dynamic AI companies (e.g., Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, Writer). By combining applied AI research, flexible infrastructure, and seamless developer tooling, we help frontier AI companies bring cutting-edge models into production. We’re growing quickly and recently raised our **$1.5B Series F** (led by Altimeter Capital, Conviction Partners, Spark Capital). Join us and help build the platform engineers rely on to ship AI products. **THE ROLE** As a **Software Engineer on the Training Infrastructure** team, you’ll architect and lead development of our training platform—supporting top-tier research engineers and model developers. You’ll make key technical decisions for infrastructure that enables developers to **deploy, scale, and monitor** training workloads with **high performance and reliability**. You’ll own **scheduling, storage, networking, reliability, and observability** across the training stack. **EXAMPLE INITIATIVES** - Product overview: https://www.baseten.co/blog/baseten-training-is-ga/#training-is-now-ga - Training docs overview: https://docs.baseten.co/training/overview - Training product story: https://www.baseten.co/blog/a-q-a-from-inference-to-training-the-inside-story-of-baseten-s-newest-product/ - Research: https://www.baseten.co/resources/research/ **RESPONSIBILITIES** - Design and architect scalable infrastructure systems for the ML training platform (e.g., **scheduling, storage, networking**) - Partner with developers and research engineers to translate complex training requirements into technical solutions - Design and architect a **global training scheduler** - Design and architect **reinforcement learning systems** and **continuous learning pipelines** - Drive long-term improvements to increase **reliability** and **development velocity** - Partner with **SRE** and **Capacity** teams to unlock state-of-the-art training infrastructure - Make critical architectural decisions balancing **performance** with **system reliability** - Lead technical discussions and mentor junior engineers on infrastructure best practices - Contribute to long-term technical strategy and infrastructure roadmap **REQUIREMENTS** - Bachelor’s degree (or higher) in Computer Science or related field - Proficiency in **Go** - Deep expertise with **Kubernetes** in production environments - Advanced understanding of **distributed systems** and **performance tuning** - Proven experience designing **observability** systems - Experience with **ML/AI workloads** and **MLOps** platforms **NICE TO HAVE** - Experience with distributed storage systems - Python experience - Extensive experience with major cloud providers (**AWS, GCP**) and neo-cloud providers (**Crusoe, DigitalOcean, Nebius**) - Experience with workload orchestration platforms like **Temporal** or **Airflow** - Familiarity with open source training stack/frameworks (**NCCL, PyTorch, Megatron, NemoRL, VeRL, Axolotl, HF Trainer**) and distributed training techniques (**FSDP, DeepSpeed**) - Experience developing AI products, tooling, or agents **BENEFITS** - Competitive compensation, including meaningful equity - (U.S. only) 100% coverage of medical, dental, and vision insurance for employees and dependents - Flexible PTO, including company-wide Winter Break (offices closed **Dec 24–Jan 1**) - Paid parental leave - Fertility and family-building stipend through Carrot - (U.S. only) C
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.