2 papers
cs.CL2026
JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment
Russell Yang, Ruishi Chen, Pierce Kelaita +6
Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise prefer…
cs.CY2025
Towards Robust Legal Reasoning: Harnessing Logical LLMs in Law
Manuj Kant, Sareh Nabi, Manav Kant +3
Legal services rely heavily on text processing. While large language models (LLMs) show promise, their application in legal contexts demands higher accuracy, repeatability, and tra…