1 paper
Shumeng Yang, Yisu Liu, Jiayi Zheng +2
Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning but often suffers from rapid policy-entropy collapse, where the policy prematurely con…