3 papers
cs.AI2026
Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners
Feng Xiong, Leyan Xue, Hongyu Lin
On-policy distillation provides dense supervision for multimodal reasoners, but its trajectory-level reward cannot determine whether a failed answer arose from perception or subseq…
cs.LG2026
Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models
Leyan Xue, Feng Xiong, Mingjun Ma +1
On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefi…
cs.CV2025
Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach
Leyan Xue, Zongbo Han, Guangyu Wang +3
Vision-Language Models (VLMs) like CLIP achieve cross-modal semantic alignment through contrastive learning, exhibiting robust zero-shot generalization. Traditional prompt engineer…