Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Towards a Theoretical Understanding to the Generalization of RLHF
Zhaochun Li, Mingyang Yi, Yue Wang +2
Reinforcement Learning from Human Feedback (RLHF) and its variants have emerged as the dominant approaches for aligning Large Language Models with human intent. While empirically e…
cs.LG2025
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
Xuerui Su, Yue Wang, Jinhua Zhu +4
With the rapid development of Large Language Models (LLMs), numerous Reinforcement Learning from Human Feedback (RLHF) algorithms have been introduced to improve model safety and a…