collaborators

5 papers

cs.CV2026

ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision

Delin Mao, Chenghao Sun, Jingwei Song +2

Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then-RL recipe creates a different…

cs.LG2026

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

Yi Yang, Cong Qin, Xiaodan Liu +8

Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens…

cs.CL2026

Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation

Chishui Chen, Yaoyou Fan, Te Sun +11

On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn age…

cs.CL2026

Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning

Chishui Chen, Jiaye Lin, Te Sun +6

Agent skills are callable procedural modules that provide reusable knowledge and execution policies for complex agentic tasks. However, existing methods mainly focus on selecting r…

cs.MA2025

Towards Robust Multi-UAV Collaboration: MARL with Noise-Resilient Communication and Attention Mechanisms

Zilin Zhao, Chishui Chen, Haotian Shi +4

Efficient path planning for unmanned aerial vehicles (UAVs) is crucial in remote sensing and information collection. As task scales expand, the cooperative deployment of multiple U…