papers

Publications (8)

cs.CR2026

BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment

Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis +13

As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refusal, a benchmark for risk iden…

cs.AI2026

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

Hannah Le, Ramesh Ramasamy, Alex Urrutia +3

Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluati…

cs.AI2026

SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?

Kenny Workman, Zhen Yang, Harihara Muralidharan +1

Spatial transcriptomics assays are rapidly increasing in scale and complexity, making computational analysis a major bottleneck in biological discovery. Although frontier AI agents…

q-bio.GN2026

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology

Ian Diks, Zhen Yang, Arjun Banerjee +2

Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxili…

cs.AI2026

Verifiable Benchmarking of Long-Horizon Spatial Biology

Ian Diks, Harihara Muralidharan, Tim Proctor +1

AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or localized analysis steps ra…

cs.AI2026

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

Harmon Bhasin, Kevin Flyangolts, Dianzhuo Wang +9

As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations…