3 papers
cs.LG2024
On Designing Effective RL Reward at Training Time for LLM Reasoning
Jiaxuan Gao, Shusheng Xu, Wenjie Ye +6
Reward models have been increasingly critical for improving the reasoning capability of LLMs. Existing research has shown that a well-trained reward model can substantially improve…
cs.DC2024
ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation
Zhiyu Mei, Wei Fu, Kaiwei Li +3
Reinforcement Learning from Human Feedback (RLHF) is a pivotal technique for empowering large language model (LLM) applications. Compared with the supervised training process of LL…
cs.CL2024
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Shusheng Xu, Wei Fu, Jiaxuan Gao +6
Reinforcement Learning from Human Feedback (RLHF) is currently the most widely used method to align large language models (LLMs) with human preferences. Existing RLHF methods can b…