3 papers
cs.LG2026
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning
Renjie Mao, Xiangxin Zhou, Lvfang Tao +7
Reinforcement learning with verifiable rewards (RLVR) has become standard for improving LLM reasoning. However, existing PPO-style trust-region mechanisms remain position-agnostic…
cs.LG2024
Entropy-Regularized Process Reward Model
Hanning Zhang, Pengcheng Wang, Shizhe Diao +6
Large language models (LLMs) have shown promise in performing complex multi-step reasoning, yet they continue to struggle with mathematical reasoning, often making systematic error…
cs.LG2024
On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
Yong Lin, Skyler Seto, Maartje ter Hoeve +6
Reinforcement Learning from Human Feedback (RLHF) is an effective approach for aligning language models to human preferences. Central to RLHF is learning a reward function for scor…