1 paper
Zelin Zhang, Fei Cheng, Chenhui Chu
Although outcome-based reinforcement learning (RL) significantly advances the mathematical reasoning capabilities of Large Language Models (LLMs), its reliance on computationally e…