Applied AI Researcher, Agent Systems & Evaluation
Nuro · Mountain View, California (HQ)
About this role
**Applied AI Researcher, Agent Systems & Evaluation** **About Nuro** Nuro is a self-driving technology company building the world's most scalable driver by combining cutting-edge AI with automotive-grade hardware. Founded in 2016, Nuro licenses its core technology to support applications from robotaxis to personally owned vehicles. **About the Team** This team builds the infrastructure that decides whether an autonomous system's output can be trusted—evaluation, verification, and evidence-based gating. We operate as a startup within a company that has already shipped a hard thing, with direct access to compute and the systems we're automating. **The Role** You'll establish rigorous, evidence-based evaluation for AI agent systems operating inside Nuro's engineering organization. Every decision about agent construction should be settled by evidence, not intuition. **Key Responsibilities:** • Own the evaluation pipeline end-to-end: data collection, loop construction, and automated hill climbing • Make the system perform against real-world data and production traffic • Post-train models on proprietary driving data using available labeling capacity • Read the research frontier and convert it into live experiments • Map test-time scaling strategies: reasoning budgets, sampling, verification, and compute allocation • Establish evaluation foundations, task suites, and statistical standards for the agent fleet **What You Bring:** • Graduate degree in CS, ML, statistics, or equivalent research experience • Fluency in current LLM literature with strong research judgment • Deep understanding of LLM architecture, pretraining, post-training, and inference • Full evaluation pipeline experience: data sourcing, loop construction, and automated optimization • Rigorous experimental design with comfort in negative results • Hands-on post-training experience (SFT, RL) on open-weight models • Strong Python and production systems experience • Direct LLM agent systems experience • Measure yourself in impact and weeks **Bonus:** Agent evaluation, reasoning, test-time compute, RL, verification, multimodal/VLM experience, online experimentation, evaluation harnesses, or autonomous systems background.
Listing freshness
CronJobs last confirmed this listing 1d ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.