From the 1 of 20 linked papers with an AI index.
7 papers · 1 filter
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes…
Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation
Jinwei Xiao, Zhuowen Han, Yueqing Sun +6
On-policy distillation transfers reasoning ability through dense token-level supervision, yet the nature of the transferable signal remains unclear. We discover that reasoning chai…
MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning
Yaorui Shi, Shugui Liu, Yu Yang +7
Long-horizon agentic reasoning necessitates effectively compressing growing interaction histories into a limited context window. Most existing memory systems serialize history as t…
MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning
Yuxin Liu, Ziang Ye, Yueqing Sun +6
Current interactive LLM agents rely on goal-conditioned stepwise planning, where environmental understanding is acquired reactively during execution rather than established beforeh…
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning
Yaorui Shi, Yuxin Chen, Zhengxi Lu +6
A persistent skill library allows language model agents to reuse successful strategies across tasks. Maintaining such a library requires three coupled capabilities. The agent selec…
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
Wentao Shi, Yu Wang, Yuyang Zhao +8
As reinforcement learning continues to scale the training of large language model-based agents, reliably verifying agent behaviors in complex environments has become increasingly c…