Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
Xilun Chen, Zhaleh Feizollahi, Ross Goodwin +5
Rubric-based evaluation of open-ended generation faces a fundamental tension between expressiveness and reliability. Authoring a faithful rubric requires expressing the structure o…
cs.CL2024
Doppelgänger's Watch: A Split Objective Approach to Large Language Models
Shervin Ghasemlou, Ashish Katiyar, Aparajita Saraf +5
In this paper, we investigate the problem of "generation supervision" in large language models, and present a novel bicameral architecture to separate supervision signals from thei…