1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Zhonghang Yuan, Zhefan Wang, Fang Hu +7
Reinforcement learning with verifiable rewards (RLVR) has demonstrated promising potential to enhance the reasoning capabilities of large language models (LLMs) in domains such as…