1 citations · 1 across the 1 of their papers we have counts for
1 paper
Jonathan D. Chang, Wenhao Zhan, Owen Oertell +4
Reinforcement Learning (RL) from Human Preference-based feedback is a popular paradigm for fine-tuning generative models, which has produced impressive models such as GPT-4 and Cla…