CronJobs

devops-sre jobs

Site Reliability Engineer, Data Center Infrastructure

SpaceX · Bastrop, TX

onsiteunknownPosted Oct 8, 2026LinuxTerraformAnsiblePuppetDockerKubernetesPostgresClickhouse

Apply on the employer site

About this role

**SpaceX — Site Reliability Engineer, Data Center Infrastructure** SpaceX is building the technologies to enable human life on Mars. The Application Software team is the “central nervous system” of SpaceX—ensuring the compute, storage, and networking that run our factories are as reliable as the products we build. This role directly impacts factory uptime, throughput, and production scale across programs. --- **Responsibilities** - Deploy, upgrade, operate, maintain, and scale compute, storage, and networking for manufacturing systems across Starship, Starlink, Starshield, and Terafab - Manage infrastructure as code and use observability to provide a complete picture of platform health - Design for reliability, stability, and scale; identify and remove bottlenecks using measurement and engineering - Practice proactive maintenance (capacity planning, lifecycle management, and reducing toil before incidents) - Partner with software engineers, manufacturing stakeholders, and site teams to build operable, maintainable systems - Improve the full lifecycle—from design through deployment, operation, and continuous refinement - Practice sustainable incident response and blameless postmortems - Provide high-quality support to manufacturing and engineering users - Communicate clearly with stakeholders and teammates - Participate in on-call and travel to sites as needed for deployments, incidents, and cross-site reliability --- **Basic Qualifications** - Bachelor’s degree in computer science, information systems, or an engineering discipline; **or** 3+ years of professional experience in SRE or DevOps in lieu of a degree - 1+ years of software development experience - Experience with Linux operating systems --- **Preferred Skills & Experience** - Experience with compute, storage, and/or networking infrastructure in production - Infrastructure as code (Terraform, Ansible, Puppet, or similar) - Containers and virtualization (Docker, Kubernetes, vSphere, QEMU, KVM, etc.) - Databases and data modeling (Postgres, Clickhouse, etc.) - Ability to translate high-level requirements into implementations from first principles - Comfort operating mission-critical systems with appropriate urgency and care - Strong communication with customers, peers, and management - Comfort operating across multiple sites and manufacturing programs --- **Additional Requirements** - Ability to work extended hours and weekends as needed - Ability to travel to sites (Hawthorne, CA; Redmond, WA; Cape Canaveral, FL; Starbase, TX) - Ability to pass an Air Force background check for Cape Canaveral - **Onsite required** — remote and/or hybrid work will not be considered --- **ITAR Requirements** To conform to U.S. Government export regulations, applicants must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (green card holder), (iii) Refugee under 8 U.S.C. § 1157, (iv) Asylee under 8 U.S.C. § 1158, or eligible to obtain required authorizations from the U.S. Department of State. --- SpaceX is an Equal Opportunity Employer.

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord