collaborators

5 papers

cs.CV2026

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models

Mengjie Zhang, Qihui Zhu, Tao Zhang +10

Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number of redundant spatio-temporal v…

cs.CV2026

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models

Qihui Zhu, Yuchen Wang, Zijian Wen +7

On-Policy Self-Distillation (OPSD) uses privileged information available only to the teacher to provide dense token-level supervision on trajectories generated by the student. Howe…

cs.CV2026

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models

Qihui Zhu, Tao Zhang, Yuchen Wang +9

In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making them impractical for real-time…

cs.CV2025

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning

Yi Liu, Xiao Xu, Zeyu Xu +10

Vision-Language Models (VLMs) have achieved remarkable breakthroughs in recent years, enabling a diverse array of applications in everyday life. However, the substantial computatio…

cs.CV2025

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Desen Meng, Rui Huang, Zhilin Dai +8

While recent advances in reinforcement learning have significantly enhanced reasoning capabilities in large language models (LLMs), these techniques remain underexplored in multi-m…