1 paper
Hao Wang, Guozhi Wang, Han Xiao +8
Reinforcement learning (RL) has been widely used to train LLM agents for multi-turn interactive tasks, but its sample efficiency is severely limited by sparse rewards and long hori…