3 papers
cs.CL2026
Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
Linlu Qiu, Fei Sha, Kelsey Allen +3
Large language models (LLMs) are increasingly used as agents that interact with users and with the world. To do so successfully, LLMs must construct representations of the world an…
cs.AI2025
Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
Nitya Thakkar, Mert Yuksekgonul, Jake Silberg +6
Peer review at AI conferences is stressed by rapidly rising submission volumes, leading to deteriorating review quality and increased author dissatisfaction. To address these issue…
cs.LG2025
Graders should cheat: privileged information enables expert-level automated evaluations
Jin Peng Zhou, Sébastien M. R. Arnold, Nan Ding +3
Auto-evaluating language models (LMs), i.e., using a grader LM to evaluate the candidate LM, is an appealing way to accelerate the evaluation process and the cost associated with i…