5 papers
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
Yuchun Miao, Sen Zhang, Yuqi Zhang +4
Reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for advanced reasoning in Large Language Models (LLMs), but rollout samples are expensive to…
Aligning Few-Step Diffusion Models with Dense Reward Difference Learning
Ziyi Zhang, Li Shen, Sen Zhang +6
Few-step diffusion models enable efficient high-resolution image synthesis but struggle to align with specific downstream objectives due to limitations of existing reinforcement le…
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
Yuchun Miao, Liang Ding, Sen Zhang +3
Despite the success of Reinforcement Learning from Human Feedback (RLHF) in aligning language models with human values, reward hacking-or reward over-optimization-remains a major c…
Image Captions are Natural Prompts for Text-to-Image Models
Shiye Lei, Hao Chen, Sen Zhang +2
With the rapid development of Artificial Intelligence Generated Content (AIGC), it has become a common practice to train models on synthetic data due to data-scarcity and privacy l…
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
Yuchun Miao, Sen Zhang, Liang Ding +3
This work identifies the Energy Loss Phenomenon in Reinforcement Learning from Human Feedback (RLHF) and its connection to reward hacking. Specifically, energy loss in the final la…