activity
20242026
collaborators

6 papers

cs.AI2026

LABBench2: An Improved Benchmark for AI Systems Performing Biology Research

Jon M Laurent, Albert Bou, Michael Pieler +9

Optimism for accelerating scientific discovery with AI continues to grow. Current applications of AI in scientific research range from training dedicated foundation models on scien…

cs.AI2025

Kosmos: An AI Scientist for Autonomous Discovery

Ludovico Mitchener, Angela Yiu, Benjamin Chang +34

Data-driven scientific discovery requires iterative cycles of literature search, hypothesis generation, and data analysis. Substantial progress has been made towards AI agents that…

q-bio.QM2025

BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology

Ludovico Mitchener, Jon M Laurent, Alex Andonian +6

Large Language Models (LLMs) and LLM-based agents show great promise in accelerating scientific research. Existing benchmarks for measuring this potential and guiding future develo…

cs.LG2025

Benchmarking and Evaluation of AI Models in Biology: Outcomes and Recommendations from the CZI Virtual Cells Workshop

Elizabeth Fahsbender, Alma Andersson, Jeremy Ash +32

Artificial intelligence holds immense promise for transforming biology, yet a lack of standardized, cross domain, benchmarks undermines our ability to build robust, trustworthy mod…

cs.AI2025

Robin: A multi-agent system for automating scientific discovery

Ali Essam Ghareeb, Benjamin Chang, Ludovico Mitchener +7

Scientific discovery is driven by the iterative process of background research, hypothesis generation, experimentation, and data analysis. Despite recent advancements in applying a…

cs.AI2024

Aviary: training language agents on challenging scientific tasks

Siddharth Narayanan, James D. Braza, Ryan-Rhys Griffiths +8

Solving complex real-world tasks requires cycles of actions and observations. This is particularly true in science, where tasks require many cycles of analysis, tool use, and exper…