3 citations · 3 across the 2 of their papers we have counts for
1 paper · 1 filter
Manon Revel, Matteo Cargnelutti, Tyna Eloundou +1
Reinforcement Learning from Human Feedback (RLHF) aims to align language models (LMs) with human values by training reward models (RMs) on binary preferences and using these RMs to…