1 paper · 1 filter
Changyi Xiao, Caijun Xu, Yixin Cao
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective in enhancing the reasoning capabilities of large language models, particularly in domains such as mathema…