9 papers
AI Alignment From Social Choice Perspectives
Daniel Halpern, Evi Micha, Ariel D. Procaccia +3
Alignment from human feedback uses human judgments about model outputs to steer the behavior of language models after pretraining. When those judgments reflect conflicting views of…
Robust AI Evaluation through Maximal Lotteries
Hadi Khalaf, Serena L. Wang, Daniel Halpern +3
The standard way to evaluate language models on subjective tasks is through pairwise comparisons: an annotator chooses the "better" of two responses to a prompt. Leaderboards aggre…
How RLHF Amplifies Sycophancy
Itai Shapira, Gerdus Benade, Ariel D. Procaccia
Large language models often exhibit increased sycophantic behavior after preference-based post-training, showing a stronger tendency to affirm a user's stated or implied belief eve…
Incentives in Federated Learning with Heterogeneous Agents
Ariel D. Procaccia, Han Shao, Itai Shapira
Federated learning promises significant sample-efficiency gains by pooling data across multiple agents, yet incentive misalignment is an obstacle: each update is costly to the cont…
Pairwise Calibrated Rewards for Pluralistic Alignment
Daniel Halpern, Evi Micha, Ariel D. Procaccia +1
Current alignment pipelines presume a single, universal notion of desirable behavior. However, human preferences often diverge across users, contexts, and cultures. As a result, di…
Generative Social Choice
Sara Fish, Paul Gölz, David C. Parkes +4
The mathematical study of voting, social choice theory, has traditionally only been applicable to choices among a few predetermined alternatives, but not to open-ended decisions su…