From the 1 of 17 linked papers with an AI index.
5 papers · 1 filter
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes…
Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation
Jinwei Xiao, Zhuowen Han, Yueqing Sun +6
On-policy distillation transfers reasoning ability through dense token-level supervision, yet the nature of the transferable signal remains unclear. We discover that reasoning chai…
MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning
Yuxin Liu, Ziang Ye, Yueqing Sun +6
Current interactive LLM agents rely on goal-conditioned stepwise planning, where environmental understanding is acquired reactively during execution rather than established beforeh…
ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training
Dunwei Tu, Hongyan Hao, Hansi Yang +10
Training generalist agents capable of adapting to diverse scenarios requires interactive environments for self-exploration. However, interactive environments remain critically scar…
Cross-Modal Memory Compression for Efficient Multi-Agent Debate
Jing Wu, Yue Sun, Tianpei Xie +7
Multi-agent debate can improve reasoning quality and reduce hallucinations, but it incurs rapidly growing context as debate rounds and agent count increase. Retaining full textual…