2 papers
cs.LG2026
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents
Edward De Brouwer, Carl Edwards, Alexander Wu +9
Recent advances in machine learning and large-scale biological data collections have revived the prospect of building a virtual cell, a computational model of cellular behavior tha…
cs.CL2026
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
Xinming Tu, Tianze Wang, Yingzhou +4
As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken specifications, implicit ass…