2 papers
cs.AI2026
Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence
Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar +2
LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy sa…
cs.LG2026
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
Stephane Collot, Colin Fraser, Justin Zhao +3
Rigorous evaluation of large language models (LLMs) relies on comparing models by the prevalence of desirable or undesirable behaviors, such as task pass rates or policy violations…