6 papers
LABBench2: An Improved Benchmark for AI Systems Performing Biology Research
Jon M Laurent, Albert Bou, Michael Pieler +9
Optimism for accelerating scientific discovery with AI continues to grow. Current applications of AI in scientific research range from training dedicated foundation models on scien…
Kosmos: An AI Scientist for Autonomous Discovery
Ludovico Mitchener, Angela Yiu, Benjamin Chang +34
Data-driven scientific discovery requires iterative cycles of literature search, hypothesis generation, and data analysis. Substantial progress has been made towards AI agents that…
BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
Ludovico Mitchener, Jon M Laurent, Alex Andonian +6
Large Language Models (LLMs) and LLM-based agents show great promise in accelerating scientific research. Existing benchmarks for measuring this potential and guiding future develo…
Benchmarking and Evaluation of AI Models in Biology: Outcomes and Recommendations from the CZI Virtual Cells Workshop
Elizabeth Fahsbender, Alma Andersson, Jeremy Ash +32
Artificial intelligence holds immense promise for transforming biology, yet a lack of standardized, cross domain, benchmarks undermines our ability to build robust, trustworthy mod…
Robin: A multi-agent system for automating scientific discovery
Ali Essam Ghareeb, Benjamin Chang, Ludovico Mitchener +7
Scientific discovery is driven by the iterative process of background research, hypothesis generation, experimentation, and data analysis. Despite recent advancements in applying a…
Aviary: training language agents on challenging scientific tasks
Siddharth Narayanan, James D. Braza, Ryan-Rhys Griffiths +8
Solving complex real-world tasks requires cycles of actions and observations. This is particularly true in science, where tasks require many cycles of analysis, tool use, and exper…