1 paper · 1 filter
Wenkai Fang, Shunyu Liu, Yang Zhou +5
Recent advances have demonstrated the effectiveness of Reinforcement Learning (RL) in improving the reasoning capabilities of Large Language Models (LLMs). However, existing works…