3 papers
cs.CL2025
PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning
Afra Feyza Akyürek, Advait Gosai, Chen Bo Calvin Zhang +21
Frontier model progress is often measured by academic benchmarks, which offer a limited view of performance in real-world professional contexts. Existing evaluations often fail to…
cs.CL2025
Online Rubrics Elicitation from Pairwise Comparisons
MohammadHossein Rezaei, Robert Vacareanu, Zihao Wang +4
Rubrics provide a flexible way to train LLMs on open-ended long-form answers where verifiable rewards are not applicable and human preferences provide coarse signals. Prior work sh…
cs.CL2024
Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
Afra Feyza Akyürek, Ekin Akyürek, Leshem Choshen +2
While language models (LMs) can sometimes generate factually correct text and estimate truth values of individual claims, these generally do not reflect a globally coherent, manipu…