CronJobs

devops-sre jobs

Network Engineer (Supercomputer Infrastructure)

xAI · Memphis, Tennessee; Southaven, Mississippi

onsitemidPosted Sep 2, 2026PythonTerraformAnsibleCiscoAristaJuniperRoCEv2GitOps

Apply on the employer site

About this role

**Network Engineer (Supercomputer Infrastructure)** **ABOUT THE ROLE:** SpaceXAI is seeking an exceptional network engineer with experience in mission-critical, large-scale production environments to support the design, build-out, and operation of networks powering AI supercomputer campuses. You'll provide design and operational support for fabrics used by GPU training and inference clusters, site operations, automation, and facilities teams. **KEY RESPONSIBILITIES:** • Design and implement highly available, low-latency, high-bandwidth networks for AI training fabrics, inference front-ends, and storage • Design and maintain supercomputer data center and campus networks per company standards • Evaluate, procure, and deploy network hardware (400G/800G switches, NICs, firewalls, optical multiplexers) • Contribute to network automation tooling; implement configuration analysis and GitOps/IaC frameworks • Coordinate network change windows and perform maintenance including evenings/weekends • Troubleshoot network issues affecting cluster health; publish RCA documentation • Provide direct support during cluster bring-up and production campaigns; serve as on-call engineer • Proactively monitor fabric health, congestion, and collective performance • Create and maintain network documentation and operational procedures • Ensure compliance with industry and cybersecurity standards (ITAR, ISO, NIST) **BASIC QUALIFICATIONS:** • Bachelor's in CS/Engineering/STEM + 3+ years network engineering experience (OR 5+ years without degree) • Extensive hands-on experience with Layer 2/3 networks in latency-sensitive/data-center environments • Functional experience with multiple network vendors in production • Experience with GitOps and Infrastructure as Code frameworks **PREFERRED SKILLS:** • Strong OSI model and network standards knowledge • Hands-on experience with Cisco, Arista, Juniper, or NVIDIA Spectrum-X switches • RoCEv2 Ethernet AI/HPC fabric experience; InfiniBand a plus • Understanding of AI training/inference traffic patterns and NCCL • WDM and large-scale fiber plant experience • Network segmentation, QoS, multicast, and redundancy protocols • Scripting proficiency (Bash/PowerShell/Python) and automation frameworks • Linux and Windows system administration experience • CCNA or CCNP certifications • Experience with real-time systems or OT networks **ADDITIONAL REQUIREMENTS:** • Ability to pass background checks for site access • Physical dexterity; ability to work in tight quarters • Availability for extended hours/weekends and 24x7 on-call support • Willingness to travel up to 20% between campuses • Ability to lift 30 lbs, work at heights, and drive (valid license) *SpaceXAI is an equal opportunity employer.*

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord