1 paper
Yun Qu, Qi Wang, Yixiu Mao +3
Recent advances have witnessed the effectiveness of reinforcement learning (RL) finetuning in enhancing the reasoning capabilities of large language models (LLMs). The optimization…