4 papers
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents
Haoze Wu, Chuqiao Kuang, Tianyi Zhuang +1
Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory, which is far too sp…
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
Junlong Li, Wenshuo Zhao, Jian Zhao +18
Real-world language agents must handle complex, multi-step workflows across diverse Apps. For instance, an agent may manage emails by coordinating with calendars and file systems,…
ReCode: Updating Code API Knowledge with Reinforcement Learning
Haoze Wu, Yunzhi Yao, Wenhao Yu +1
Large Language Models (LLMs) exhibit remarkable code generation capabilities but falter when adapting to frequent updates in external library APIs. This critical limitation, stemmi…
Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
Haoze Wu, Cheng Wang, Wenshuo Zhao +1
Recent advances in applying reinforcement learning (RL) to large language models (LLMs) have led to substantial progress. In particular, a series of remarkable yet often counterint…