2 papers
cs.LG2026
Towards a Theoretical Understanding to the Generalization of RLHF
Zhaochun Li, Mingyang Yi, Yue Wang +2
Reinforcement Learning from Human Feedback (RLHF) and its variants have emerged as the dominant approaches for aligning Large Language Models with human intent. While empirically e…
cs.LG2026
Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off
Zhaochun Li, Chen Wang, Jionghao Bai +4
The exploration-exploitation (EE) trade-off is a central challenge in reinforcement learning (RL) for large language models (LLMs). With Group Relative Policy Optimization (GRPO),…