5 papers
Adaptive Contracts for Cost-Effective AI Delegation
Eden Saig, Tamar Garbuz, Ariel D. Procaccia +2
When organizations delegate text generation tasks to AI providers via pay-for-performance contracts, expected payments rise when evaluation is noisy. As evaluation methods become m…
Pluralistic Leaderboards
Nika Haghtalab, Ariel D. Procaccia, Han Shao +2
Recent leaderboard-based evaluations of large language models aggregate user feedback by fitting a Bradley--Terry model to pairwise comparisons, producing a single global ranking b…
Robust AI Evaluation through Maximal Lotteries
Hadi Khalaf, Serena L. Wang, Daniel Halpern +3
The standard way to evaluate language models on subjective tasks is through pairwise comparisons: an annotator chooses the "better" of two responses to a prompt. Leaderboards aggre…
The Hidden Cost of Waiting for Accurate Predictions
Ali Shirali, Ariel Procaccia, Rediet Abebe
Algorithmic predictions are increasingly informing societal resource allocations by identifying individuals for targeting. Policymakers often build these systems with the assumptio…
Direct Alignment with Heterogeneous Preferences
Ali Shirali, Arash Nasr-Esfahany, Abdullah Alomar +3
Alignment with human preferences is commonly framed using a universal reward function, even though human preferences are inherently heterogeneous. We formalize this heterogeneity b…