activity
20242026
collaborators

11 papers

cs.AI2026

How RLHF Amplifies Sycophancy

Itai Shapira, Gerdus Benade, Ariel D. Procaccia

Large language models often exhibit increased sycophantic behavior after preference-based post-training, showing a stronger tendency to affirm a user's stated or implied belief eve…

cs.GT2026

Incentives in Federated Learning with Heterogeneous Agents

Ariel D. Procaccia, Han Shao, Itai Shapira

Federated learning promises significant sample-efficiency gains by pooling data across multiple agents, yet incentive misalignment is an obstacle: each update is costly to the cont…

cs.AI2025

Question the Questions: Auditing Representation in Online Deliberative Processes

Soham De, Lodewijk Gelauff, Ashish Goel +3

A central feature of many deliberative processes, such as citizens' assemblies and deliberative polls, is the opportunity for participants to engage directly with experts. While pa…

cs.LG2025

Alternates, Assemble! Selecting Optimal Alternates for Citizens' Assemblies

Angelos Assos, Carmel Baharav, Bailey Flanigan +1

Citizens' assemblies are an increasingly influential form of deliberative democracy, where randomly selected people discuss policy questions. The legitimacy of these assemblies hin…

cs.LG2025

Metritocracy: Representative Metrics for Lite Benchmarks

Ariel Procaccia, Benjamin Schiffer, Serena Wang +1

A common problem in LLM evaluation is how to choose a subset of metrics from a full suite of possible metrics. Subset selection is usually done for efficiency or interpretability r…

cs.LG2025

Pairwise Calibrated Rewards for Pluralistic Alignment

Daniel Halpern, Evi Micha, Ariel D. Procaccia +1

Current alignment pipelines presume a single, universal notion of desirable behavior. However, human preferences often diverge across users, contexts, and cultures. As a result, di…