From the 1 of 6 linked papers with an AI index.
1 paper · 1 filter
Wenbin Wu
Reinforcement Learning from Human Feedback (RLHF) assumes annotator preferences reflect stable internal states. We challenge this through three experiments spanning the preference…