3 papers
cs.CL2026
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment
Yu Li, Xiuyu Li, Mingyang Yi +4
Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of "rollout then update", which inevitably results in off-policy training da…
cs.LG2026
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
Kaibing Yang, Guangfeng Cai, Shengtian Yang +6
Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group.…
cs.CL2026
Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making
Guangfeng Cai, Kaibing Yang, Shuo He +4
Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties super…