2 papers
cs.CL2024
PERSONA: A Reproducible Testbed for Pluralistic Alignment
Louis Castricato, Nathan Lile, Rafael Rafailov +2
The rapid advancement of language models (LMs) necessitates robust alignment with diverse user values. However, current preference optimization approaches often fail to capture the…
cs.CL2023
Social Contract AI: Aligning AI Assistants with Implicit Group Norms
Jan-Philipp Fränken, Sam Kwok, Peixuan Ye +6
We explore the idea of aligning an AI assistant by inverting a model of users' (unknown) preferences from observed interactions. To validate our proposal, we run proof-of-concept s…