2 papers
cs.CL2026
Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias
Justin D. Norman, Michael U. Rivera, D. Alex Hughes
LLM-as-a-Judge has become the dominant evaluation paradigm for language models, but judge validation in practice relies on exact-match agreement, a metric that does not correct for…
cs.CL2025
The Case for Repeatable, Open, and Expert-Grounded Hallucination Benchmarks in Large Language Models
Justin D. Norman, Michael U. Rivera, D. Alex Hughes
Plausible, but inaccurate, tokens in model-generated text are widely believed to be pervasive and problematic for the responsible adoption of language models. Despite this concern,…