Site Reliability Engineer, AI Infrastructure (Starshield)
SpaceX · Redmond, WA
About this role
🚀 **SpaceX — Site Reliability Engineer (AI Infrastructure / Starshield)** SpaceX is developing technologies to enable human life on Mars. Starshield is the world’s largest US government satellite constellation, providing immediate access to critical intelligence and national security data for the US government anywhere on the globe. As a **Site Reliability Engineer focused on Starshield’s software and GPU infrastructure**, you’ll design, operate, and scale infrastructure supporting critical national security missions. This role spans areas including **Site Reliability Engineering, Developer Operations, and GPU platforms**—with opportunities to build automation, create scalable software products, and collaborate closely with engineering teams. --- ## Responsibilities - Manage **GPU/CPU infrastructure deployments** to **Top Secret datacenters** - Manage and support **GPU as a service** for external customers on **bare metal** and **virtualized** platforms - Design, validate, and productize solutions for **AI clusters (100k+ GPU scale)** - Develop automation to deploy and manage **on-premise Kubernetes/AI clusters** and **operating systems** - Deploy and manage core infrastructure such as **databases, monitoring, and distributed storage** - Collaborate with AI engineers to build **highly scalable, operable, and maintainable** products - Improve services across the full lifecycle: **inception → design → deployment → operation → refinement** - Build **monitoring and alerting** to support **high availability** - Identify improvement opportunities and create innovative solutions to increase system availability --- ## Basic Qualifications - Bachelor’s degree in computer science, information systems/IT, or an engineering discipline **and 1+ years** of professional experience in **SRE or DevOps**; **or** **3+ years** of professional experience in SRE/DevOps in lieu of a degree - **1+ years** of professional experience with **Linux** - Experience with **Terraform, Ansible, or other infrastructure tools** - Experience with **containerization technologies** (e.g., **OCI containers, Kubernetes**) - Experience scripting in **Bash, Python, or other** languages
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.