1 paper · 1 filter
Chenyi Li, Yuan Zhang, Bo Wang +4
Reinforcement learning with verifiable rewards has shown notable effectiveness in enhancing large language models (LLMs) reasoning performance, especially in mathematics tasks. How…