11 citations · 11 across the 1 of their papers we have counts for
1 paper
Yao Zhao, Rishabh Joshi, Tianqi Liu +3
Learning from human feedback has been shown to be effective at aligning language models with human preferences. Past work has often relied on Reinforcement Learning from Human Feed…