1 paper
Hao Zhang, Mingjie Liu, Shaokun Zhang +10
Multi-turn LLM agents are increasingly important for solving complex, interactive tasks, and reinforcement learning (RL) is a key ingredient for improving their long-horizon behavi…