1 paper · 1 filter
Yulun Jiang, Liangze Jiang, Damien Teney +2
Reinforcement learning (RL) has enabled the training of large language model (LLM) agents to interact with the environment and to solve multi-turn long-horizon tasks. However, the…