activity
20242026
collaborators

6 papers

cs.LG2026

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning

Renjie Mao, Xiangxin Zhou, Lvfang Tao +7

Reinforcement learning with verifiable rewards (RLVR) has become standard for improving LLM reasoning. However, existing PPO-style trust-region mechanisms remain position-agnostic…

cs.LG2025

Entropy-Regularized Process Reward Model

Hanning Zhang, Pengcheng Wang, Shizhe Diao +6

Large language models (LLMs) have shown promise in performing complex multi-step reasoning, yet they continue to struggle with mathematical reasoning, often making systematic error…

cs.CL2024

Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Rui Yang, Ruomeng Ding, Yong Lin +2

Reward models trained on human preference data have been proven to effectively align Large Language Models (LLMs) with human intent within the framework of reinforcement learning f…

cs.LG2024

Mitigating the Alignment Tax of RLHF

Yong Lin, Hangyu Lin, Wei Xiong +14

LLMs acquire a wide range of abilities during pre-training, but aligning LLMs under Reinforcement Learning with Human Feedback (RLHF) can lead to forgetting pretrained abilities, w…

cs.LG2024

On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization

Yong Lin, Skyler Seto, Maartje ter Hoeve +6

Reinforcement Learning from Human Feedback (RLHF) is an effective approach for aligning language models to human preferences. Central to RLHF is learning a reward function for scor…

cs.CV2024

The Instinctive Bias: Spurious Images lead to Illusion in MLLMs

Tianyang Han, Qing Lian, Rui Pan +5

Large language models (LLMs) have recently experienced remarkable progress, where the advent of multi-modal large language models (MLLMs) has endowed LLMs with visual capabilities,…