1 paper · 1 filter
Xucong Wang, Ziyu Ma, Yong Wang +5
Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods of…