1 paper
Yifan Xu, Junren Chen, Yifan Chen
Reinforcement learning with verifiable rewards (RLVR) recently thrives in large language model (LLM) reasoning tasks. However, the reward sparsity and the long reasoning horizon ma…