activity
20242026
collaborators

13 papers

cs.CV2026

ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs

Yuhao Wang, Mu Qiao, Haiwen Diao +5

Multimodal Large Language Models (MLLMs) incur prohibitive inference costs due to long visual token sequences. Training-free visual token reduction provides an efficient solution.…

cs.CV2026

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

Qing'an Liu, Juntong Feng, Yuhao Wang +6

Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing benchmarks predominantly focus on pur…

cs.CV2025

Regularizing Subspace Redundancy of Low-Rank Adaptation

Yue Zhu, Haiwen Diao, Shang Gao +8

Low-Rank Adaptation (LoRA) and its variants have delivered strong capability in Parameter-Efficient Transfer Learning (PETL) by minimizing trainable parameters and benefiting from…

cs.CV2025

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Haiwen Diao, Xiaotong Li, Yufeng Cui +6

Existing encoder-free vision-language models (VLMs) are rapidly narrowing the performance gap with their encoder-based counterparts, highlighting the promising potential for unifie…

cs.CV2025

End-to-End Vision Tokenizer Tuning

Wenxuan Wang, Fan Zhang, Yufeng Cui +5

Existing vision tokenization isolates the optimization of vision tokenizers from downstream training, implicitly assuming the visual tokens can generalize well across various tasks…

cs.CV2025

Exploring Dynamic Transformer for Efficient Object Tracking

Jiawen Zhu, Xin Chen, Haiwen Diao +6

The speed-precision trade-off is a critical problem for visual object tracking which usually requires low latency and deployment on constrained resources. Existing solutions for ef…