1 paper · 1 filter
Yue Cheng, Jiajun Zhang, Xiaohui Gao +3
Reinforcement Learning with Verifiable Reward (RLVR) is empirically shown to notably enhance the reasoning performance of large language models (LLMs), particularly in mathematics…