4.3k citations · 5.8k across the 18 of their papers we have counts for
1 paper · 2 filters
Leo Gao, John Schulman, Jacob Hilton
In reinforcement learning from human feedback, it is common to optimize against a reward model trained to predict human preferences. Because the reward model is an imperfect proxy,…