Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets
Zongyou Yang, Yinghan Hou
Language-model benchmarks collapse two distinct measurement questions into a single accuracy score: whether a response reached an evaluable state, and whether its answer was judged…
cs.CL2026
When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability
Zongyou Yang, Yinghan Hou, Xiaokun Yang
An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed. We treat this evaluator-replacement ambiguity as a measuremen…