1 paper
Haotian Zhao, Songlin Zhou, Yuxin Zhang +9
Reinforcement learning (RL) has substantially improved the ability of large language model (LLM) agents to interact with environments and solve multi-turn tasks. However, effective…