2 citations · 2 across the 1 of their papers we have counts for
1 paper
Michael Santacroce, Yadong Lu, Han Yu +2
Reinforcement Learning with Human Feedback (RLHF) has revolutionized language modeling by aligning models with human preferences. However, the RL stage, Proximal Policy Optimizatio…