1 paper
Zixuan Wang, Yuchen Yan, Hongxing Li +7
While long-horizon agentic tasks require language agents to perform dozens of sequential decisions, training such agents with reinforcement learning remains challenging. We identif…