CronJobs

devops-sre jobs

Site Reliability Engineer (High Performance Computing)

SpaceX · Hawthorne, CA

remoteunknown$145,000–$195,000Posted Sep 23, 2026LinuxPythonKubernetesTerraformAnsiblePuppetDockerPrometheus

Apply on the employer site

About this role

**SpaceX — Site Reliability Engineer (High Performance Computing)** SpaceX is building the technologies to enable human life on Mars. SpaceX HPC is a shared compute platform used across the company for vehicle and structures simulation, machine learning, AI inference, and more. This role puts a real SRE operating model behind these capabilities—accelerating world-class engineering through toil reduction, automation, observability, and a sustainable incident process. You’ll own the full ecosystem as a product: Linux machines, Infrastructure as Code, storage, and user-facing applications. You don’t need prior HPC experience, but you do need production instincts—operating real infrastructure, writing code to delete toil, and caring about whether users can actually get work done. You’ll work alongside HPC systems engineers who design and commission clusters to improve reliability and deliver world-class services. --- **Responsibilities** - Participate in the team’s on-call rotation; practice sustainable incident response and blameless postmortems - Manage node lifecycle with Infrastructure as Code (OS images, firmware, configuration management, kernel/driver stack) - Build observability for both HPC administrators and end users (cluster/node/storage health + job/workflow-level signal) - Reduce toil with automation; split time between operating production systems and building software that makes operations smaller - Sustainably manage resources, including compute and storage - Lead capacity planning with users across the company; translate needs into a concrete picture of tomorrow’s compute and storage - Collaborate with HPC systems engineers and engineers across disciplines to deliver operable, maintainable infrastructure --- **Basic Qualifications** - Bachelor’s degree in CS/engineering/math/science **or** 2+ years of professional experience operating production infrastructure in lieu of a degree - 2+ years of experience with Linux operating systems in production - 2+ years of experience operating production infrastructure (servers/services/networks), including monitoring, debugging, and repairing what you own --- **Preferred Skills & Experience** - 2+ years in SRE, DevOps, or production infrastructure engineering - Monitoring/alerting experience (Prometheus, Grafana, Nagios, or similar) - Configuration management and/or Infrastructure as Code (Ansible, Puppet, Terraform, or similar) - Scripting/coding to automate common tasks (e.g., Python) - Containers (Docker, Podman, Singularity/Apptainer) - Kubernetes administration for on-prem deployments - Distributed/high-performance storage experience (e.g., VAST) including capacity/performance/lifecycle management - Familiarity with HPC clusters/schedulers (Slurm, PBS, LSF) or GPU compute (not required; training provided) - Familiarity with scientific computing (CFD/FEA) and/or ML training (PyTorch/TensorFlow/CUDA) and/or AI inference workloads - Strong understanding of version control, testing, CI/CD, build/deployment, and monitoring - Clear communication with users, peers, and vendors in incident and design settings - Comfortable working with mission-critical/sensitive systems with appropriate urgency - Eligibility for access to classified material up to TS/SCI with polygraph --- **Additional Requirements** - Based in Hawthorne, CA; primarily on-site - Must participate in an on-call rotation - Willing to work extended hours and weekends as needed for incidents, cluster bring-up, and

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord