1 paper · 1 filter
Ziyang Luo, Yan Yang, Xiangru Jian +5
Reinforcement learning (RL) has become an effective way to improve the tool-use ability of large language models (LLMs), but most existing RL frameworks stop at the policy update.…