Compute Systems Architect - Expeditionary AI/HPC Data Centers
Anduril Industries · Costa Mesa, California, United States
About this role
**Anduril Industries — Compute Systems Architect (Expeditionary AI/HPC Data Centers)** **About the Job** Anduril Expeditionary Systems (AES) is building AI Factories and Command/Control centers—specialized, high-performance computing systems designed to process petabyte-scale data and convert it into decisions at machine speed from forward operating bases. These are **air-transportable, power-constrained, hardened** systems optimized for **AI inference at the tactical edge** and deeply integrated with **Lattice OS**. This is a **new, highest-priority role**. You will architect the complete electrical and physical infrastructure—**compute node selection, power distribution, thermal management, battery backup, and physical hardening**—to enable AI processing in **austere, contested, and DDIL (degraded/denied/intermittent/limited)** environments. --- **What You’ll Do** - Own the **end-to-end architecture** of AI Factory tactical data centers (compute hardware → power architecture → networking → thermal cooling) - Design **GPU-accelerated compute clusters** optimized for AI inference/training at the tactical edge (e.g., **NVIDIA H100/A100, AMD MI300, or similar**) - Engineer **advanced thermal solutions** for high-density compute without traditional data center infrastructure (liquid, immersion, hybrid) - Design **air-transportable, modular systems** (containerized, palletized, vehicle-mounted) that are rapidly deployable - Build infrastructure optimized for **extreme temperatures, limited grid power, SCIF-accreditable, and electromagnetic hardening** - Work closely with **Lattice** and **AI/ML teams** to ensure physical infrastructure meets mission-critical software requirements - Partner with **defense customers** to translate operational needs into technical architectures - Lead **prototyping, field testing, and operational validation** through exercises and real-world deployments - Drive systems from **concept through production at scale** across hardware/software/supply chain/program teams - Define **standards and best practices** for tactical AI compute infrastructure across AES product lines --- **Required Qualifications** - Bachelor’s degree in Computer Engineering, Computer Science, Electrical Engineering, or related field - **5+ years** designing/deploying data center infrastructure, HPC clusters, large-scale compute systems, or tensor/accelerator compute nodes - **8+ years** engineering experience designing/deploying electrical products - Deep expertise in **GPU-accelerated computing architecture** for AI/ML (NVIDIA/AMD or similar) - Strong understanding of **power distribution** for high-density compute (PDUs, power shelves, UPS, generators, chillers, power distribution, energy management) - Experience with **thermal management** for high-power-density systems (liquid cooling/chiller systems) - Proven track record building/deploying production infrastructure from concept through operational scale - Systems thinking across competing constraints (performance, power, thermal, weight, cost, reliability) - **U.S. Person status required** (U.S. citizen, permanent resident, refugee, or asylee) - Ability to obtain/maintain a **U.S. security clearance** (U.S. citizenship may be required) - Ability to travel up to **15%** --- **Preferred Qualifications** - Background in AI/ML infrastructure (large-scale inference, computer vision, autonomous systems) - Experience with high-speed compute networking and network topologies - Familiar
Listing freshness
CronJobs last confirmed this listing 1d ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.