Publications (4)
TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology
Hannah Le, Ramesh Ramasamy, Alex Urrutia +3
Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluati…
scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology
Ian Diks, Zhen Yang, Arjun Banerjee +2
Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxili…
Verifiable Benchmarking of Long-Horizon Spatial Biology
Ian Diks, Harihara Muralidharan, Tim Proctor +1
AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or localized analysis steps ra…
EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis
Harihara Muralidharan, Reema Baskar, Soo Hee Lee +2
We introduce EpiBench, a verifiable benchmark for short-horizon epigenomics analysis. EpiBench evaluates whether agents can make well-defined analysis decisions from realistic work…