2 papers
cs.AI2026
COMPOSITE-Stem
Kyle Waters, Lucas Nuzzi, Tadhg Looram +20
AI agents hold growing promise for accelerating scientific discovery; yet, a lack of frontier evaluations hinders adoption into real workflows. Expert-written benchmarks have prove…
cs.LG2026
AsymmetryZero: A Framework for Operationalizing Human Expert Preferences as Semantic Evals
Tadhg Looram, Lucas Nuzzi, Kyle Waters +1
Much of the focus in RL today is on evaluation design: building meaningful evals that serve simultaneously as benchmarks and as well-defined reward signals for post-training. Yet,…