1 paper
Yulei Qin, Xiaoyu Tan, Zhengbao He +13
Reinforcement learning (RL) is the dominant paradigm for sharpening strategic tool use capabilities of LLMs on long-horizon, sparsely-rewarded agent tasks, yet it faces a fundament…