3 papers
cs.LG2025
Process Reinforcement through Implicit Rewards
Ganqu Cui, Lifan Yuan, Zefan Wang +22
Dense process rewards have proven a more effective alternative to the sparse outcome-level rewards in the inference-time scaling of large language models (LLMs), particularly in ta…
cs.LG2025
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Ganqu Cui, Yuchen Zhang, Jiacheng Chen +14
This paper aims to overcome a major obstacle in scaling RL for reasoning with LLMs, namely the collapse of policy entropy. Such phenomenon is consistently observed across vast RL r…
cs.LG2024
Free Process Rewards without Process Labels
Lifan Yuan, Wendi Li, Huayu Chen +6
Different from its counterpart outcome reward models (ORMs), which evaluate the entire responses, a process reward model (PRM) scores a reasoning trajectory step by step, providing…