2 citations · 2 across the 2 of their papers we have counts for
3 papers · 1 filter
Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment
Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez +1
When an AI algorithm makes decisions that affect more than one person, aligning it becomes a problem of social choice: how should people's divergent preferences about system behavi…
How RLHF Amplifies Sycophancy
Itai Shapira, Gerdus Benade, Ariel D. Procaccia
Large language models often exhibit increased sycophantic behavior after preference-based post-training, showing a stronger tendency to affirm a user's stated or implied belief eve…
Question the Questions: Auditing Representation in Online Deliberative Processes
Soham De, Lodewijk Gelauff, Ashish Goel +3
A central feature of many deliberative processes, such as citizens' assemblies and deliberative polls, is the opportunity for participants to engage directly with experts. While pa…