2 papers
cs.CL2026
JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment
Russell Yang, Ruishi Chen, Pierce Kelaita +6
Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise prefer…
cs.CL2026
Reflections and New Directions for Human-Centered Large Language Models
Caleb Ziems, Dora Zhao, Rose E. Wang +55
Large Language Models (LLMs) are increasingly shaping the private and professional lives of users, with numerous applications in business, education, finance, healthcare, law, and…