4 papers
Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence
Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar +2
LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy sa…
Muse Spark Safety & Preparedness Report
Cristina Menghini, Peter Ney, Hamza Kwisaba +117
Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
Stephane Collot, Colin Fraser, Justin Zhao +3
Rigorous evaluation of large language models (LLMs) relies on comparing models by the prevalence of desirable or undesirable behaviors, such as task pass rates or policy violations…
Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks
Justin Zhao, Flor Miriam Plaza-del-Arco, Benjamin Genchel +1
As Large Language Models (LLMs) continue to evolve, evaluating them remains a persistent challenge. Many recent evaluations use LLMs as judges to score outputs from other LLMs, oft…