2 citations · 2 across the 2 of their papers we have counts for
6 papers
How RLHF Amplifies Sycophancy
Itai Shapira, Gerdus Benade, Ariel D. Procaccia
Large language models often exhibit increased sycophantic behavior after preference-based post-training, showing a stronger tendency to affirm a user's stated or implied belief eve…
Incentives in Federated Learning with Heterogeneous Agents
Ariel D. Procaccia, Han Shao, Itai Shapira
Federated learning promises significant sample-efficiency gains by pooling data across multiple agents, yet incentive misalignment is an obstacle: each update is costly to the cont…
Metritocracy: Representative Metrics for Lite Benchmarks
Ariel Procaccia, Benjamin Schiffer, Serena Wang +1
A common problem in LLM evaluation is how to choose a subset of metrics from a full suite of possible metrics. Subset selection is usually done for efficiency or interpretability r…
Alternates, Assemble! Selecting Optimal Alternates for Citizens' Assemblies
Angelos Assos, Carmel Baharav, Bailey Flanigan +1
Citizens' assemblies are an increasingly influential form of deliberative democracy, where randomly selected people discuss policy questions. The legitimacy of these assemblies hin…
Pairwise Calibrated Rewards for Pluralistic Alignment
Daniel Halpern, Evi Micha, Ariel D. Procaccia +1
Current alignment pipelines presume a single, universal notion of desirable behavior. However, human preferences often diverge across users, contexts, and cultures. As a result, di…
The Proportional Veto Principle for Approval Ballots
Daniel Halpern, Ariel D. Procaccia, Warut Suksompong
The proportional veto principle, which captures the idea that a candidate vetoed by a large group of voters should not be chosen, has been studied for ranked ballots in single-winn…