13 papers
WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training
Zehao Chen, Gongxun Li, Tianxiang Ai +9
On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state mismatch of offline distillation. The sa…
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
Zixuan Huang, Yang Zhou, Kaixuan Wang +7
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision…
Policy Improvement Reinforcement Learning
Huaiyang Wang, Xiaojie Li, Xiaohan Wang +10
Reinforcement learning has become a central post-training paradigm for improving LLM and agent capabilities. Yet existing RL post-training methods share a common blind spot: they c…
Multi-Objective Exploration and Preference Optimization via Mutual Information
Hongyan Xie, Yikun Ban, Ruiyu Fang +4
Aligning large language models with diverse and heterogeneous human values requires multi-objective alignment methods to effectively trade off conflicting preference dimensions. Cu…
Weak-Driven Learning: How Weak Agents make Strong Agents Stronger
Zehao Chen, Gongxun Li, Tianxiang Ai +9
As post-training optimization becomes central to improving large language models, we observe a persistent saturation bottleneck: once models grow highly confident, further training…
Heterogeneous Agent Collaborative Reinforcement Learning
Zhixia Zhang, Zixuan Huang, Gongxun Li +10
We introduce Heterogeneous Agent Collaborative Reinforcement Learning (HACRL), a new Reinforcement Learning from Verifiable Reward (RLVR) problem that addresses the inefficiencies…