4 papers
Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence
Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar +2
LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy sa…
Muse Spark Safety & Preparedness Report
Cristina Menghini, Peter Ney, Hamza Kwisaba +117
Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…
SCRuB: Social Concept Reasoning under Rubric-Based Evaluation
Jamelle Watson-Daniels, Himaghna Bhattacharjee, Skyler Wang +11
While many studies of Large Language Model (LLM) reasoning capabilities emphasize mathematical or technical tasks, few address reasoning about social concepts: the abstract ideas s…
Counterfactual Evaluation of Ads Ranking Models through Domain Adaptation
Mohamed A. Radwan, Himaghna Bhattacharjee, Quinn Lanners +4
We propose a domain-adapted reward model that works alongside an Offline A/B testing system for evaluating ranking models. This approach effectively measures reward for ranking mod…