2 papers
cs.CL2026
JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment
Russell Yang, Ruishi Chen, Pierce Kelaita +6
Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise prefer…
cs.AI2026
RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics
Zhengyang Qi, Charles Dickens, Derek Pham +4
Rubric-based evaluation is widely used in LLM benchmarks and training pipelines for open-ended, less verifiable tasks. While prior work has demonstrated the effectiveness of rubric…