Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Aligning Language Models with Demonstrated Feedback
Omar Shaikh, Michelle S. Lam, Joey Hejna +4
Language models are aligned to emulate the collective voice of many, resulting in outputs that align with no one in particular. Steering LLMs away from generic output is possible t…
cs.CL2024
PERSONA: A Reproducible Testbed for Pluralistic Alignment
Louis Castricato, Nathan Lile, Rafael Rafailov +2
The rapid advancement of language models (LMs) necessitates robust alignment with diverse user values. However, current preference optimization approaches often fail to capture the…