1 paper · 1 filter
Han Zhong, Zikang Shan, Guhao Feng +6
In the classical Reinforcement Learning from Human Feedback (RLHF) framework, Proximal Policy Optimization (PPO) is employed to learn from sparse, sentence-level rewards -- a chall…