collaborators

5 papers

cs.AI2026

When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation

Yinuo Jiang, Yongjie Ye, Zhou Tao +4

On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve…

cs.CV2026

LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models

Zhou Tao, Fang Zhang, Zewen Ding +5

Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessary local details. We identify thi…

cs.CV2026

Dynamic Token Compression for Efficient Video Understanding through Reinforcement Learning

Shida Wang, YongXiang Hua, Zhou Tao +2

Multimodal Large Language Models have demonstrated remarkable capabilities in video understanding, yet face prohibitive computational costs and performance degradation from ''conte…

cs.CV2026

DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model

Zhou Tao, Shida Wang, Yongxiang Hua +2

Multimodal Large Language Models have achieved impressive performance on a variety of vision-language tasks, yet their fine-grained visual perception and precise spatial reasoning…

cs.LG2025

Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts

Yongxiang Hua, Haoyu Cao, Zhou Tao +4

Sparse Mixture of Experts (sMoE) has become a pivotal approach for scaling large vision-language models, offering substantial capacity while maintaining computational efficiency th…