Site Reliability Engineer, Cyber
Anduril Industries · Arlington, Virginia, United States
About this role
**Anduril Industries — Site Reliability Engineer (Cyber)** ## About Anduril Anduril is a defense technology company on a mission to transform U.S. and allied military capabilities with advanced technology. Anduril’s family of systems is powered by **Lattice OS**, an AI-powered operating system that turns thousands of data streams into a real-time, 3D command and control center. ## About the Team (Anduril Cyber) Anduril Cyber focuses on positioning Anduril as a lead provider of capabilities to enable **offensive cyber missions**. This new and fast-growing business line relies on autonomous vehicles, the Lattice operating system, mesh networks, and other hardware products to deploy cyber capabilities at the edge—often in unconventional or difficult-to-reach environments. ## About the Job As a **Site Reliability Engineer** in Anduril Cyber, you’ll solve problems across **networking, systems integration, and distributed systems**, making pragmatic engineering tradeoffs to ensure software is **reliable, scalable, and deployable**. You’ll own the full deployment pipeline—from **CI/CD and pre-production test environments**, through **canary deployments** in customer-hosted integration environments, to **production deployments in air-gapped enclaves**. You’ll also act as a technical steward with customers: attending meetings, explaining system behavior, absorbing requirements firsthand, and feeding onsite observations into the team’s roadmap. ## What You’ll Do - Own the health of deployed systems and keep them running with minimal downtime - Automate and improve software deployment processes into **air-gapped, TS/SCI** environments - Design, build, and maintain **CI/CD** and automated test infrastructure for complex hardware/software systems - Develop metrics dashboards, TUIs, scripts, and tools to automate deployment steps and debug the software stack - Drive engineering requirements based on onsite observations - Perform root cause analysis across the software stack, **Lattice OS**, and external vendor services - Build strong relationships with internal and external customers to identify technical solutions - Lead continuous improvement by instrumenting systems, analyzing failures, and running post-mortems spanning software, firmware, and hardware ## Required Qualifications - Active U.S. **TS/SCI** security clearance (and ability to maintain it) - Based in the **DC metro area** to support **3–5 days/week** on site at customer facilities - **4+ years** in Sys Admin, Site Reliability, DevOps, or Software Engineering - Deep, practical experience with **Linux** and **Kubernetes** (or similar container orchestration) - Working knowledge of network fundamentals and ability to debug connectivity in locked-down environments - Experience delivering and maintaining systems on **air-gapped and security-hardened** networks - Strong proficiency in **Python or Bash** for automation/debugging; ability to read/debug service code in a compiled language such as **Go** - Excellent written and verbal communication skills for cross-functional collaboration and customer engagement ## Preferred Qualifications - Experience debugging and resolving networking issues - Ability to quickly understand and navigate complex, multi-disciplinary systems and established codebases - Experience building automation for **HIL** or **SIL** test environments - Ability to drive consensus across internal and external stakeholders ## Salary Range **$146,000 – $220,000 USD** (b
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.