1 citations · 4 across the 34 of their papers we have counts for
10 papers · 1 filter
One Step, One Lead: Mitigating Higher-Order Interference in Multi-Domain Reinforcement Learning via Cross-Step Control
Zihan Lin, Xiaohan Wang, Jie Cao +4
Reinforcement learning (RL) across multiple domains can broaden the reasoning capabilities of large language models (LLMs), yet joint training often degrades individual-domain perf…
Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning
Li Wang, Xiaodong Lu, Xiaohan Wang +4
Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already pres…
Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration
Zili Wang, Jiajun Chai, Lin Chen +3
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as the standard paradigm for improving reasoning capability of large language models, while Multi-Token Prediction…
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards
Li Wang, Xiaodong Lu, Xiaohan Wang +5
Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewards (RLVR). Nonetheless, RLVR i…
-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data
Yaocheng Zhang, Yuanheng Zhu, Wenyue Chong +7
Deep search agents have emerged as a promising paradigm for addressing complex information-seeking tasks, but their training remains challenging due to sparse rewards, weak credit…
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
Zihan Lin, Xiaohan Wang, Jie Cao +6
Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diversity due to the over-incentivi…