works on

From the 1 of 17 linked papers with an AI index.

collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10

Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes…

cs.AI2026

Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation

Jinwei Xiao, Zhuowen Han, Yueqing Sun +6

On-policy distillation transfers reasoning ability through dense token-level supervision, yet the nature of the transferable signal remains unclear. We discover that reasoning chai…

cs.AI2026

MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning

Yuxin Liu, Ziang Ye, Yueqing Sun +6

Current interactive LLM agents rely on goal-conditioned stepwise planning, where environmental understanding is acquired reactively during execution rather than established beforeh…

cs.AI2026

ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training

Dunwei Tu, Hongyan Hao, Hansi Yang +10

Training generalist agents capable of adapting to diverse scenarios requires interactive environments for self-exploration. However, interactive environments remain critically scar…

cs.AI2026

Cross-Modal Memory Compression for Efficient Multi-Agent Debate

Jing Wu, Yue Sun, Tianpei Xie +7

Multi-agent debate can improve reasoning quality and reduce hallucinations, but it incurs rapidly growing context as debate rounds and agent count increase. Retaining full textual…