collaborators

6 papers

cs.LG2026

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

Kaibing Yang, Guangfeng Cai, Shengtian Yang +6

Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group.…

cs.CL2026

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment

Yu Li, Xiuyu Li, Mingyang Yi +5

Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of "rollout then update", which inevitably results in off-policy training da…

cs.CL2026

Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making

Guangfeng Cai, Kaibing Yang, Shuo He +4

Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties super…

cs.AI2026

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning

Yu Li, Mingyang Yi, Xiuyu Li +6

Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most existing ARL methods train a sin…

cs.AI2026

Phase-Aware Mixture of Experts for Agentic Reinforcement Learning

Shengtian Yang, Yu Li, Shuo He +4

Reinforcement learning (RL) has equipped LLM agents with a strong ability to solve complex tasks. However, existing RL methods normally use a \emph{single} policy network, causing…

cs.LG2026

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment

Xiuyu Li, Jinkai Zhang, Mingyang Yi +4

Reinforcement Learning (RL) post-training alignment for language models is effective, but also costly and unstable in practice, owing to its complicated training process. To addres…