4 papers
VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation
Zewen Ding, Zezhong Wu, Zhou Tao +5
On-policy self-distillation (OPSD) improves reasoning by training a problem-only student on its own rollouts using dense token-level supervision from a privileged teacher that also…
LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models
Zhou Tao, Fang Zhang, Zewen Ding +5
Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessary local details. We identify thi…
Dynamic Token Compression for Efficient Video Understanding through Reinforcement Learning
Shida Wang, YongXiang Hua, Zhou Tao +2
Multimodal Large Language Models have demonstrated remarkable capabilities in video understanding, yet face prohibitive computational costs and performance degradation from ''conte…
DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
Zhou Tao, Shida Wang, Yongxiang Hua +2
Multimodal Large Language Models have achieved impressive performance on a variety of vision-language tasks, yet their fine-grained visual perception and precise spatial reasoning…