1 citations · 1 across the 5 of their papers we have counts for
6 papers
Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning
Li Wang, Xiaodong Lu, Xiaohan Wang +4
Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already pres…
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards
Li Wang, Xiaodong Lu, Xiaohan Wang +5
Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewards (RLVR). Nonetheless, RLVR i…
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
Xiaodong Lu, Xiaohan Wang, Jiajun Chai +7
Reinforcement Learning with Verifiable Rewards (RLVR) is an effective paradigm for improving the reasoning capabilities of large language models. However, existing RLVR methods uti…
Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning
Li Wang, Xiaohan Wang, Xiaodong Lu +5
Large language models (LLMs) have increasingly leveraged tool invocation to enhance their reasoning capabilities. However, existing approaches typically tightly couple tool invocat…
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
Zihan Lin, Xiaohan Wang, Jie Cao +6
Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diversity due to the over-incentivi…
Your Group-Relative Advantage Is Biased
Fengkai Yang, Zherui Chen, Xiaohan Wang +10
Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such…