Staff Software Engineer - AI Compute, Together Cloud
Together AI · San Francisco
About this role
**Staff Software Engineer - AI Compute, Together Cloud** **About the Role** Together AI is building the AI Native Cloud platform for the full generative AI lifecycle. The Together Cloud team builds GPU Clusters—a flagship IaaS product providing high-performance, AI-ready GPU clusters through a self-serve console, plus the virtualized infrastructure powering inference, RL, and fine-tuning products. As a Staff Software Engineer, you'll set technical direction for major components of the next-generation AI cloud platform—a highly available, global infrastructure with cutting-edge virtualization of latest ML hardware (GB300s/VRs, BlueField DPUs, InfiniBand, RoCEv2 fabrics). This architect-and-build role spans from greenfield data center IaaS layers to global management planes scheduling capacity across dozens of data centers and hundreds of thousands of GPUs. **Key Responsibilities** • Own GPU and network virtualization stack (hypervisor, kernel, SDN) • Architect in-DC IaaS layer services, Kubernetes operators, and libraries • Design GPU scheduling and global management plane for distributed clusters • Architect monitoring and automated remediation for fault tolerance • Set technical direction across teams and lead design reviews • Mentor senior and junior engineers; raise hiring bar • Create testing frameworks and developer documentation **Requirements** • 7+ years professional software development; expert-level backend language proficiency (Golang preferred) • Track record owning architecture of large distributed systems from conception to production scale • Deep experience with globally distributed, high-performance microservice architectures (AWS/Azure/GCP) • Expert systems knowledge: compute, networking, storage, concurrency, memory management, I/O • Demonstrated technical leadership: mentoring, design reviews, cross-team alignment • Excellent communication and diplomacy skills • Experience building reliable, customer-facing production systems with infrastructure automation, observability, CI/CD **Preferred Qualifications** Kubernetes internals, VMs/hypervisors, DC networking, Cluster API, high-performance compute, GPU/InfiniBand virtualization, IaaS/PaaS systems, DPUs, GPU programming (NCCL, CUDA) **Compensation** US base salary: $260,000 - $300,000 + equity + benefits. Competitive compensation, health insurance, remote work flexibility. **About Together AI** Research-driven AI company focused on lowering the cost of modern AI systems through co-designed software, hardware, algorithms, and models. Contributors to FlashAttention, Hyena, FlexGen, RedPajama, and leading open-source research. Equal Opportunity Employer. See privacy policy: https://www.together.ai/privacy
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.