3 papers
cs.CL2026
Reference-Free Evaluation of Reasoning in Open-Ended Question Answering
Guneet Singh Kohli, Yuxiang Zhou, Michael Sejr Schlichtkrull +2
AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propos…
cs.CL2026
Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
Guneet Kohli
LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a framework to measure the true inf…
cs.CL2026
Metric-Dependent Annotation Saturation for Learning from Label Distributions
Guneet Kohli
When annotators disagree on a label, the disagreement itself carries signal -- and the number of annotators needed to capture it depends on the evaluation metric. We fine-tune NLI…