4 papers
Architecting Trust in Artificial Epistemic Agents
Nahema Marchal, Stephanie Chan, Matija Franklin +5
Large language models increasingly function as epistemic agents -- entities that can 1) autonomously pursue epistemic goals and 2) actively shape our shared knowledge environment.…
AI-Enhanced Deliberative Democracy and the Future of the Collective Will
Manon Revel, Théophile Pénigaud
This article unpacks the design choices behind longstanding and newly proposed computational frameworks aimed at finding common grounds across collective preferences and examines t…
Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks
Bhaktipriya Radharapu, Manon Revel, Megan Ung +2
The increasing use of LLMs as substitutes for humans in ``aligning'' LLMs has raised questions about their ability to replicate human judgments and preferences, especially in ambiv…
SEAL: Systematic Error Analysis for Value ALignment
Manon Revel, Matteo Cargnelutti, Tyna Eloundou +1
Reinforcement Learning from Human Feedback (RLHF) aims to align language models (LMs) with human values by training reward models (RMs) on binary preferences and using these RMs to…