Staff AI Test and Evaluation Engineer
Anduril Industries · Washington, District of Columbia, United States
About this role
**Staff AI Test and Evaluation Engineer** **About Anduril** Anduril Industries is a defense technology company transforming U.S. and allied military capabilities with advanced technology. Anduril’s AI-powered operating system, **Lattice OS**, turns thousands of data streams into a real-time, 3D command and control center. **About the Team (Discovery)** Discovery is Anduril’s team for taking the newest problems across domains—space, missile systems, air, sensor capability, autonomy, and cyber—and proving out what is worth solving. They build the models, run the tests, and carry what works to the point where a program can pick it up. **About the Job** Discovery is building an expeditionary force: engineers who solve difficult problems and navigate unfamiliar territory as an operating standard. You’ll take unproven concepts and build the analysis or prototype to test them—especially as Anduril moves AI models from research into **production on classified platforms and edge hardware**. --- ## What You’ll Do - **Develop test scenarios & simulation environments** to assess agentic AI performance in classified simulations. - **Validate AI on classified & edge platforms** by leading integration and validation of agentic AI systems. - **Build reusable evaluation pipelines** (automated test harnesses, monitoring dashboards) for continuous validation. - **Define actionable metrics** for AI and traditional models, with systematic historic performance capture (not one-shot measurements). - **Partner cross-functionally** to define requirements, document test results, and contribute to AI model cards and deployment readiness reviews. - **Own end-to-end evaluation infrastructure** for agentic AI/ML models—from unit-level checks to full-system scenario-based assessment. - **Troubleshoot & debug** issues found in evaluation and deployment to ensure reliability across releases. --- ## Required Qualifications - **BS/MS** in Computer Science, Machine Learning, Software Engineering, Computer Engineering, Electrical Engineering, or related field. - **12+ years** hands-on experience writing production-grade code; strong **Python** for evaluation pipelines, test harnesses, and tooling. - Experience **owning evaluation/automated testing** for production ML or software systems (test scenarios, metrics, regression suites, CI infrastructure). - Hands-on experience designing **AI/ML evaluation methodologies** (metrics, benchmarks, real-world scenario assessment). - Experience designing and/or working extensively with **simulation environments**. - Strong **metrics discipline** (define, capture, and reason about performance over time). - Ability to navigate **complex systems and established codebases**. - Comfort operating between **technical program management and software engineering**; coordinate across teams and document results. - Passion for building evaluation infrastructure that proves AI works and impacts mission-critical outcomes. - Must be **eligible for a U.S. security clearance**. --- ## Preferred Qualifications - Background as an **ML Test & Evaluation Engineer**, SDET for ML, ML/Evaluation Engineer, Simulation Engineer, or Technical Program Manager for AI/ML. - Experience with **agentic AI evaluation** (tasking, decision-making, scenario-based behavior validation). - Familiarity with **classified/edge deployment** and validating models on constrained environments. - Experience contributing to **model cards** and **deployment readiness reviews*
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.