CronJobs

devops-sre jobs

Senior AI Infrastructure Engineer, Physical Infrastructure

Anduril Industries · Costa Mesa, California, United States

remotesenior$166,000–$166,000Posted Sep 8, 2026KubernetesNCCLRayRun:AINVLinkInfiniBandRoCENVLink

Apply on the employer site

About this role

**Senior AI Infrastructure Engineer, Physical Infrastructure** **About Anduril** Anduril Industries is a defense technology company transforming U.S. and allied military capabilities with advanced technology. Anduril’s family of systems is powered by **Lattice OS**, an AI-powered operating system that turns thousands of data streams into a real-time, 3D command and control center. **About the Team (CorpTech Infrastructure Engineering)** CorpTech Infrastructure Engineering builds and operates the foundational infrastructure that powers Anduril at large—enabling engineers, researchers, and product teams to deploy fast, scalable infrastructure without needing to become infrastructure experts. As Anduril’s AI and autonomy ambitions grow, the team delivers the next generation of **compute, networking, and storage** capabilities to support large-scale model training and inference. **About the Job** Anduril is seeking a **Senior AI Infrastructure Engineer** to lead the vision, execution, and long-term stability of how Anduril trains with GPUs at scale. You will take ownership of **cluster robustness**, ensuring high-performance GPU systems are highly available, fault-tolerant, and resilient for ML platform and research teams across the company. This is a hands-on role focused on **logical stability** and **automated resilience/self-healing** mechanisms to detect and isolate hardware faults, **tune NCCL and high-speed networking**, and optimize **Kubernetes**, **Run:AI**, and **Ray** scheduling. You’ll replace manual triage with automated deployment tooling and deep observability so massive-scale training infrastructure runs reliably and scales without linear headcount growth. **What You’ll Do** - Rack, stack, cable, and bring up GPU compute (H200/B200/B300, NVL72), including physical topology, power, cooling, firmware/BIOS, and burn-in validation. - Build and tune the interconnect fabric (NVLink, InfiniBand, RoCE, Spectrum-X) connecting hundreds of GPUs into low-latency training and inference clusters. - Integrate high-performance parallel storage (VAST, DDN, Weka) to sustain distributed training throughput and terabyte-scale multimodal datasets. - Automate cluster deployment and configuration end to end (infrastructure as code for bring-up, firmware/driver management, and fabric configuration). - Operate and extend Kubernetes/Run:AI for GPU scheduling, quota management, and multi-tenant workload isolation. - Own fleet health: monitoring, alerting, and rapid triage of hardware and network faults (e.g., bad transceivers, GPU Xid errors, NCCL/collective failures, RoCE congestion). - Onboard engineers and researchers onto the platform and serve as their escalation point to debug and optimize workloads when infrastructure is the bottleneck. - Partner with product-facing teams to understand emerging compute needs and translate them into platform capabilities. **Required Qualifications** - 10+ years in hands-on infrastructure, HPC, or datacenter engineering supporting GPU compute at scale. - Hands-on experience with H200/B200/B300 (or comparable) GPU systems: bring-up, cabling, firmware/driver management. - Experience with high-performance interconnects (NVLink, InfiniBand, RoCE, Spectrum-X) in clusters of hundreds of GPUs. - Experience with high-performance parallel storage (VAST, DDN, Weka, Lustre, or similar). - Kubernetes required; Run:AI (or similar) strongly preferred. - Strong automation background (build repeatable automated deployment

Listing freshness

CronJobs last confirmed this listing 1d ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord