1 paper
Aubrey Brueckner, Darshil Patel, Yuhuan He +1
Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a kno…