1 citations · 1 across the 10 of their papers we have counts for
5 papers · 1 filter
TREK: Distill to Explore, Reinforce to Refine
Yuanda Xu, Zhengze Zhou, Kayhan Behdin +10
Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning trajectories, but it stalls on hard prompts whose correct solution m…
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
Yuanda Xu, Hejian Sang, Zhengze Zhou +3
In settings where labeled verifiable training data is the binding constraint, each checked example should be allocated to the model and reward density where it is most informative.…
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
Zihan Lin, Xiaohan Wang, Jie Cao +6
Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diversity due to the over-incentivi…
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
Yongcan Yu, Lingxiao He, Jian Liang +5
Test-time reinforcement learning (TTRL) always adapts models at inference time via pseudo-labeling, leaving it vulnerable to spurious optimization signals from label noise. Through…
Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning
Yuanda Xu, Hejian Sang, Zhengze Zhou +2
Reinforcement Learning with Verifiable Rewards (RLVR) has become the leading paradigm for enhancing reasoning in Large Language Models (LLMs). However, standard RLVR algorithms suf…