1 paper
Yan Sun, Jia Guo, Stanley Kok +3
Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning ability of large language models, yet training remains costly because many rollouts contribute litt…