From the 1 of 4 linked papers with an AI index.
4 papers
Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models
Leyan Xue, Feng Xiong, Mingjun Ma +1
On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefi…
Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners
Feng Xiong, Leyan Xue, Hongyu Lin
The paper proposes Perception-Correction Distillation (PCD), a label‑free method that uses downstream failures and teacher‑student disagreement to pinpoint and correct perception e…
MULTIBENCH++: A Unified and Comprehensive Multimodal Fusion Benchmarking Across Specialized Domains
Leyan Xue, Changqing Zhang, Kecheng Xue +3
Although multimodal fusion has made significant progress, its advancement is severely hindered by the lack of adequate evaluation benchmarks. Current fusion methods are typically e…
Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach
Leyan Xue, Zongbo Han, Guangyu Wang +3
Vision-Language Models (VLMs) like CLIP achieve cross-modal semantic alignment through contrastive learning, exhibiting robust zero-shot generalization. Traditional prompt engineer…