4.3k citations · 4.8k across the 10 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.LG2022★ 37 cited
Scaling Laws for Reward Model Overoptimization
Leo Gao, John Schulman, Jacob Hilton
In reinforcement learning from human feedback, it is common to optimize against a reward model trained to predict human preferences. Because the reward model is an imperfect proxy,…
cs.CL2022★ 4.3k cited
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang +17
Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic,…