Software Engineer, Compute Infrastructure
OpenAI · San Francisco
About this role
**Software Engineer, Compute Infrastructure** **About the Team** Compute Infrastructure builds the platform powering frontier AI by designing, provisioning, scheduling, and optimizing systems that connect accelerators, CPUs, networks, storage, and orchestration software into a reliable engine for research and products. **About the Role** We're hiring engineers to build the compute platform behind OpenAI's research and products. You'll work across the full stack—from capacity planning and bare-metal automation to distributed systems, Kubernetes scheduling, system optimization, networking, storage, fleet health, and developer experience. **Where You Might Work** • **Compute Foundations** – Low-level platform primitives for heterogeneous hardware at scale • **Fleet / Orchestration** – Reliable, efficient clusters and scheduling systems • **Core Network Engineering** – High-performance networking fabrics and protocols • **Hardware Health & Observability** – Detect and prevent fleet-health issues • **Storage** – Scalable, performant storage abstractions • **Agent Infrastructure** – Sandboxed execution for agentic workloads **Key Responsibilities** • Build and optimize reliable system software for large-scale AI compute • Design infrastructure across accelerators, networking, storage, and orchestration • Profile and optimize training workloads across compute and networking bottlenecks • Create hardware-aware automation for provisioning and operations • Build tools and platforms that help teams launch and debug workloads **You Might Thrive If You** • Have built or operated distributed systems, HPC environments, Kubernetes clusters, or production infrastructure • Enjoy working across the stack (software, hardware, networking, systems performance) • Care about making complex infrastructure understandable and usable • Can diagnose hard problems under operational pressure • Are motivated by scale, efficiency, reliability, and measurement **Qualifications** • Strong software engineering skills and production infrastructure experience • Experience in distributed systems, operating systems, networking, RDMA, NCCL, storage, Kubernetes, scheduling, observability, GPU infrastructure, or related areas • Ability to debug complex system behavior across layers • Comfort with ambiguity and strong ownership mindset
Listing freshness
CronJobs last confirmed this listing 17h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.