Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
Can Xie, Ruotong Pan, Xiangyu Wu +4
Reinforcement Learning with Verifiable Rewards (RLVR) has shown significant promise for enhancing the reasoning capabilities of large language models (LLMs). However, prevailing al…
cs.AI2025
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning
Xiao Hu, Xingyu Lu, Liyuan Mao +6
Reinforcement learning (RL) has played an important role in improving the reasoning ability of large language models (LLMs). Some studies apply RL directly to \textit{smaller} base…