Research Engineer, Model Evaluation and Improvement
Benchling · San Francisco, CA
About this role
**Research Engineer, Model Evaluation and Improvement** **About the Role** Join Benchling's team focused on making frontier AI models better at science. You'll build datasets, evaluations, and systems that help close the gap between LLM capabilities and real-world scientific problems. Work at the intersection of software engineering, biology, and frontier AI. **Key Responsibilities** • Build datasets for evaluating and improving frontier models, turning complex scientific data into high-quality tasks for LLMs • Analyze model failure modes and run experiments across frontier models to identify improvement opportunities • Build scalable data infrastructure with pipelines that curate, transform, and validate scientific data • Collaborate with frontier AI labs on developing and evaluating new approaches for scientific tasks • Work closely with scientists to translate expert judgment into reliable evaluation criteria **Required Qualifications** • 2+ years at the intersection of biology and AI • Experience evaluating and improving scientific models or LLMs for biological applications • Experience building with LLMs and understanding their capabilities and limitations • Comfort with ambiguous problems in a rapidly evolving field • Collaborative mindset and ability to work with engineers, scientists, and research partners • Thrives in fast-paced environments with shifting priorities **Location** In-person in San Francisco, Monday-Friday **About Benchling** Benchling is the AI platform for biotech R&D, trusted by 200,000+ scientists worldwide including Sanofi, Moderna, and more than half of the world's top 50 biopharma companies. Benchling is an equal opportunity employer committed to building a diverse and inclusive team.
Listing freshness
CronJobs last confirmed this listing 17h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.