4 citations · 4 across the 6 of their papers we have counts for
8 papers
Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence
Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar +2
LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy sa…
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
Bhaktipriya Radharapu, Eshika Saxena, Kenneth Li +3
As LLM-based judges become integral to industry applications, obtaining well-calibrated uncertainty estimates efficiently has become critical for production deployment. However, ex…
Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset
Lily Hong Zhang, Smitha Milli, Karen Jusko +12
How can large language models (LLMs) serve users with varying preferences that may conflict across cultural, political, or other dimensions? To advance this challenge, this paper e…
Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks
Bhaktipriya Radharapu, Manon Revel, Megan Ung +2
The increasing use of LLMs as substitutes for humans in ``aligning'' LLMs has raised questions about their ability to replicate human judgments and preferences, especially in ambiv…
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams +99
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…
Chained Tuning Leads to Biased Forgetting
Megan Ung, Alicia Sun, Samuel J. Bell +3
Large language models (LLMs) are often fine-tuned for use on downstream tasks, though this can degrade capabilities learned during previous training. This phenomenon, often referre…