11 papers
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
Kaibing Yang, Guangfeng Cai, Shengtian Yang +6
Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group.…
Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making
Guangfeng Cai, Kaibing Yang, Shuo He +4
Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties super…
Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning
Xin Cheng, Shuo He, Lang Feng +4
Group-based reinforcement learning (RL) methods have achieved remarkable success in improving the performance of large language models (LLMs) and have been rapidly extended to agen…
Towards Safer Large Reasoning Models by Promoting Safety Decision-Making before Chain-of-Thought Generation
Jianan Chen, Zhifang Zhang, Shuo He +3
Large reasoning models (LRMs) achieved remarkable performance via chain-of-thought (CoT), but recent studies showed that such enhanced reasoning capabilities are at the expense of…
Test-Time Attention Purification for Backdoored Large Vision Language Models
Zhifang Zhang, Bojun Yang, Shuo He +5
Despite the strong multimodal performance, large vision-language models (LVLMs) are vulnerable during fine-tuning to backdoor attacks, where adversaries insert trigger-embedded sam…
Online Causal Kalman Filtering for Stable and Effective Policy Optimization
Shuo He, Lang Feng, Xin Cheng +2
Reinforcement learning for large language models suffers from high-variance token-level importance sampling (IS) ratios, which would destabilize policy optimization at scale. To im…