10 papers
scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology
Ian Diks, Zhen Yang, Arjun Banerjee +2
Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxili…
TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology
Hannah Le, Ramesh Ramasamy, Alex Urrutia +3
Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluati…
EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis
Harihara Muralidharan, Reema Baskar, Soo Hee Lee +2
We introduce EpiBench, a verifiable benchmark for short-horizon epigenomics analysis. EpiBench evaluates whether agents can make well-defined analysis decisions from realistic work…
Verifiable Benchmarking of Long-Horizon Spatial Biology
Ian Diks, Harihara Muralidharan, Tim Proctor +1
AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or localized analysis steps ra…
Scalable linearized gate set tomography
Ashe Miller, Corey Ostrove, Jordan Hines +4
Characterizing errors on many-qubit quantum computers remains a key challenge to understanding and improving the performance of these devices. Current characterization methods eith…
Simulating Quantum Error Correction beyond Pauli Stochastic Errors
Jordan Hines, Corey Ostrove, Kenneth Rudinger +4
Quantum error correction (QEC), the lynchpin of fault-tolerant quantum computing (FTQC), is designed and validated against well-behaved Pauli stochastic error models. But in real-w…