collaborators

27 papers

cs.LG2026

WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training

Zehao Chen, Gongxun Li, Tianxiang Ai +9

On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state mismatch of offline distillation. The sa…

cs.LG2026

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation

Zikun Qu, Min Zhang, Mingze Kong +5

On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other pla…

cs.AI2026

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

Zixuan Huang, Yang Zhou, Kaixuan Wang +7

Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision…

cs.LG2026

MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer

Xuefei Wang, Jialu Wang, Fengbo Zhang +6

Multi-agent systems (MAS) powered by large language models (LLMs) have emerged as a powerful paradigm for complex problem solving, where performance critically depends on the under…

cs.LG2026

Policy Improvement Reinforcement Learning

Huaiyang Wang, Xiaojie Li, Xiaohan Wang +10

Reinforcement learning has become a central post-training paradigm for improving LLM and agent capabilities. Yet existing RL post-training methods share a common blind spot: they c…

cs.CL2026

Multi-Objective Exploration and Preference Optimization via Mutual Information

Hongyan Xie, Yikun Ban, Ruiyu Fang +4

Aligning large language models with diverse and heterogeneous human values requires multi-objective alignment methods to effectively trade off conflicting preference dimensions. Cu…