5 citations · 5 across the 7 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Understanding Diversity Collapse in RLVR via the Lens of Overtraining
Suqin Yuan, Jinkun Chen, Jiyang Zheng +6
Reinforcement learning with verifiable rewards (RLVR) has become a key approach for enhancing the reasoning abilities of large language models. However, RLVR often suffers from \em…
cs.LG2026
Mitigating Mismatch within Reference-based Preference Optimization
Suqin Yuan, Xingrui Yu, Jiyang Zheng +4
Direct Preference Optimization (DPO) has become the de facto standard for offline preference alignment of large language models, but its reliance on a reference policy introduces a…
cs.LG2026
Unifying Stable Optimization and Reference Regularization in RLHF
Li He, Qiang Qu, He Zhao +4
Reinforcement Learning from Human Feedback (RLHF) has advanced alignment capabilities significantly but remains hindered by two core challenges: \textbf{reward hacking} and \textbf…