1 paper
Zhong Guan, Likang Wu, Hongke Zhao +2
Many existing studies have achieved significant improvements in the reasoning capabilities of large language models (LLMs) through reinforcement learning with verifiable rewards (R…