1 paper · 1 filter
Yuhao Wang, Xiaopeng Li, Cheng Gong +4
Reinforcement learning with verifiable rewards (RLVR) has been shown to enhance the reasoning capabilities of large language models (LLMs), enabling the development of large reason…