most citedResT: Reshaping Token-Level Policy Gradients for Tool-Use Large Language Models

1 citations · 3 across the 28 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning

Li Wang, Xiaodong Lu, Xiaohan Wang +4

Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already pres…

cs.LG2026

Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration

Zili Wang, Jiajun Chai, Lin Chen +3

Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as the standard paradigm for improving reasoning capability of large language models, while Multi-Token Prediction…

cs.LG2026

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards

Li Wang, Xiaodong Lu, Xiaohan Wang +5

Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewards (RLVR). Nonetheless, RLVR i…

cs.LG2026

-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data

Yaocheng Zhang, Yuanheng Zhu, Wenyue Chong +7

Deep search agents have emerged as a promising paradigm for addressing complex information-seeking tasks, but their training remains challenging due to sparse rewards, weak credit…

cs.LG20261 cited

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

Xiaodong Lu, Xiaohan Wang, Jiajun Chai +7

Reinforcement Learning with Verifiable Rewards (RLVR) is an effective paradigm for improving the reasoning capabilities of large language models. However, existing RLVR methods uti…

cs.LG2026

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning

Zihan Lin, Xiaohan Wang, Jie Cao +6

Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diversity due to the over-incentivi…