1 paper
Yiyang Jin, Kunzhao Xu, Hang Li +4
Reinforcement learning with verifiable rewards (RLVR) has advanced the reasoning capabilities of large language models. However, existing methods rely solely on outcome rewards, wi…