Director, Site Operations
xAI · Memphis, TN
About this role
**SpaceXAI — Director, Site Operations** **About the Role** As the Director of Site Operations, you’ll own node and rack uptime for SpaceXAI’s AI supercompute cluster—the most advanced of its kind. This role is the extreme owner of cluster health and customer Service Level Agreements across 5+ sites operating 24/7. You’ll lead a 250+ person organization of site managers, shift supervisors, and technicians, plus the site reliability engineering team that monitors cluster health and drives fault mitigation at scale. We’re looking for a hands-on operations leader who can build a culture of excellence and accountability, partner tightly across the company, and keep uptime exceptional as we grow. **Responsibilities** - **Own Cluster Uptime:** Extreme ownership of node, rack, and cluster health across 5+ sites running 24/7; accountable for customer SLAs and consistently exceptional uptime. - **Lead a Large Operations Organization:** Direct a 250+ person team across four 24/7 shifts; build a culture of excellence and accountability at every layer. - **Drive Node and Rack Remediation:** Ensure systematic recovery of failed nodes and racks through command-line and physical intervention; drive mean time to repair to the feasible minimum. - **Partner Across Functions:** Coordinate with facilities operations (power/cooling, proactive maintenance), network engineering (cluster upgrades), and tenant representatives (remediation + planned/unplanned downtime). - **Own Vendor Execution:** Direct vendors through hardware rework and field operations so repairs, replacements, and capacity work happen at the speed the cluster requires. - **Lead Site Reliability Engineering:** Own SRE for proactive cluster health monitoring, reactive fault mitigation at scale, root cause analyses, and site-wide reliability procedures and fault documentation. - **Run Data-Driven Improvement:** Use operational data to balance resources and improve uptime, repair time, and SLA performance across sites. - **Command Incidents at Scale:** Set the standard for incident response during cluster-impacting events—clear direction, fast recovery, tight communication. - **Scale Operations:** Standardize best practices across sites and grow the organization with cluster expansion. **Basic Qualifications** - Bachelor’s degree and **7+ years** of experience in large-scale operations, including **5+ years leading people leaders of technical teams** - *or* **10+ years** of experience in large-scale operations, including **5+ years leading people leaders of technical teams** **Preferred Skills and Experience** - Proven ability to lead large, multi-site, 24/7 operations organizations in fast-paced, high-responsibility settings. - Deep expertise in server hardware, cluster reliability, and data center technologies (deployment through lifecycle management). - Experience supporting compute-heavy environments (AI/ML/HPC) at scale. - Track record owning uptime, SLAs, or reliability metrics for large compute clusters. - Experience leading SRE or equivalent reliability-focused teams (root cause analysis + procedure ownership). - Strong analytical skills; ability to explain technical concepts clearly to diverse audiences (technicians through executives/tenants). - Experience partnering with vendors at scale to reduce MTTR and scale operations across multiple sites. - Familiarity with tooling/automation (e.g., **Jira, Python, Bash**) used to monitor cluster health and improve efficiency. - Enthusia
Listing freshness
CronJobs last confirmed this listing 59m ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.