1 paper
Allison Lau, Younwoo Choi, Vahid Balazadeh +3
Reinforcement Learning from Human Feedback (RLHF) is widely used to align Language Models (LMs) with human preferences. However, existing approaches often neglect individual user p…