most citedHow RLHF Amplifies Sycophancy

2 citations · 2 across the 2 of their papers we have counts for

collaborators

6 papers

cs.AI20262 cited

How RLHF Amplifies Sycophancy

Itai Shapira, Gerdus Benade, Ariel D. Procaccia

Large language models often exhibit increased sycophantic behavior after preference-based post-training, showing a stronger tendency to affirm a user's stated or implied belief eve…

cs.GT2025

Incentives in Federated Learning with Heterogeneous Agents

Ariel D. Procaccia, Han Shao, Itai Shapira

Federated learning promises significant sample-efficiency gains by pooling data across multiple agents, yet incentive misalignment is an obstacle: each update is costly to the cont…

cs.LG2025

Metritocracy: Representative Metrics for Lite Benchmarks

Ariel Procaccia, Benjamin Schiffer, Serena Wang +1

A common problem in LLM evaluation is how to choose a subset of metrics from a full suite of possible metrics. Subset selection is usually done for efficiency or interpretability r…

cs.LG2025

Alternates, Assemble! Selecting Optimal Alternates for Citizens' Assemblies

Angelos Assos, Carmel Baharav, Bailey Flanigan +1

Citizens' assemblies are an increasingly influential form of deliberative democracy, where randomly selected people discuss policy questions. The legitimacy of these assemblies hin…

cs.LG2025

Pairwise Calibrated Rewards for Pluralistic Alignment

Daniel Halpern, Evi Micha, Ariel D. Procaccia +1

Current alignment pipelines presume a single, universal notion of desirable behavior. However, human preferences often diverge across users, contexts, and cultures. As a result, di…

cs.GT2025

The Proportional Veto Principle for Approval Ballots

Daniel Halpern, Ariel D. Procaccia, Warut Suksompong

The proportional veto principle, which captures the idea that a candidate vetoed by a large group of voters should not be chosen, has been studied for ranked ballots in single-winn…