4 papers
Scaling Small Agents Through Strategy Auctions
Lisa Alazraki, William F. Shen, Yoram Bachrach +1
Small language models are increasingly viewed as a promising, cost-effective approach to agentic AI, with proponents claiming they are sufficiently capable for agentic workflows. H…
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
William F. Shen, Xinchi Qiu, Chenxi Whitehouse +6
Recently, rubrics have been used to guide LLM judges in capturing subjective, nuanced, multi-dimensional human preferences, and have been extended from evaluation to reward signals…
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
Stephane Collot, Colin Fraser, Justin Zhao +3
Rigorous evaluation of large language models (LLMs) relies on comparing models by the prevalence of desirable or undesirable behaviors, such as task pass rates or policy violations…
Training AI Co-Scientists Using Rubric Rewards
Shashwat Goel, Rishi Hazra, Dulhan Jayalath +8
AI co-scientists are emerging as a tool to assist human researchers in achieving their research goals. A crucial feature of these AI co-scientists is the ability to generate a rese…