works on

From the 1 of 17 linked papers with an AI index.

collaborators

17 papers

cs.AI2026

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10

Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes…

cs.CL2026

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance

Zhuowen Han, Jinwei Xiao, Zhengxi Lu +9

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO)…

cs.LG2026

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Zhiyuan Yao, Yuxin Chen, Zhengxi Lu +13

SkillRise introduces a reinforcement‑learning framework that lets large language model agents learn and reuse transferable skills across related tasks by curating a skill document…

cs.CL2026

CAST: Game Solvers as Turn-Level Teachers for LLM Agents

Yu Wang, Yi-Kai Zhang, Wentao Shi +8

Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR)…

cs.AI2026

Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation

Jinwei Xiao, Zhuowen Han, Yueqing Sun +6

On-policy distillation transfers reasoning ability through dense token-level supervision, yet the nature of the transferable signal remains unclear. We discover that reasoning chai…

cs.CL2026

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

Tianjie Ju, Yueqing Sun, Zheng Wu +7

Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to sustain exploration in dynamic op…