3 papers
cs.AI2026
PAEC: Position-Aware Entropy Calibration for LLM Reasoning in RLVR
Shumeng Yang, Yisu Liu, Jiayi Zheng +2
Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning but often suffers from rapid policy-entropy collapse, where the policy prematurely con…
cs.AI2025
Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning
Zhaohui Yang, Yuxiao Ye, Shilei Jiang +4
Recent advances in reasoning language models have witnessed a paradigm shift from short to long CoT pattern. Given the substantial computational cost of rollouts in long CoT models…
cs.AI2025
Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning
Zhaohui Yang, Chenghua He, Xiaowen Shi +4
Many studies focus on data annotation techniques for training effective PRMs. However, current methods encounter a significant issue when applied to long CoT reasoning processes: t…