From the 2 of 45 linked papers with an AI index.
45 papers
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
Zishan Xu, Zhiyuan Yao, Yuxin Chen +9
Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verificatio…
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
Shuai Shao, Kangning Zhang, Qingyao Li +7
Agents built around large language models continually accumulate interaction trajectories during deployment, yet their behavior typically remains fixed. Beyond updating model weigh…
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
Kangning Zhang, Yixing Li, Shuai Shao +9
The paper proposes Visual Attribution Distillation (VAD), a counterfactual method that isolates the visual component of teacher corrections in multimodal on‑policy distillation and…
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
Zhiyuan Yao, Yuxin Chen, Zhengxi Lu +13
SkillRise introduces a reinforcement‑learning framework that lets large language model agents learn and reuse transferable skills across related tasks by curating a skill document…
BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation
Ning Li, Zixuan Guo, Yan Xu +7
Hallucinations remain a major obstacle to deploying large language models (LLMs) in knowledge-intensive settings, where generated responses must be faithfully grounded in provided…
Communication Policy Evolution for Proactive LLM Agents
Xinbei Ma, Jiyang Qiu, Yao Yao +10
LLM agents have rapidly evolved into autonomous systems, yet a persistent information gap remains between users and agents: communication is costly, while users' identical preferen…