1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2024
PERSONA: A Reproducible Testbed for Pluralistic Alignment
Louis Castricato, Nathan Lile, Rafael Rafailov +2
The rapid advancement of language models (LMs) necessitates robust alignment with diverse user values. However, current preference optimization approaches often fail to capture the…
cs.CL2024★ 1 cited
Aligning Language Models with Demonstrated Feedback
Omar Shaikh, Michelle S. Lam, Joey Hejna +4
Language models are aligned to emulate the collective voice of many, resulting in outputs that align with no one in particular. Steering LLMs away from generic output is possible t…