24 citations · 35 across the 14 of their papers we have counts for
10 papers · 1 filter
ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs
Yuhao Wang, Mu Qiao, Haiwen Diao +5
Multimodal Large Language Models (MLLMs) incur prohibitive inference costs due to long visual token sequences. Training-free visual token reduction provides an efficient solution.…
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
Yi Zhong, Haotong Qin, Xindong Zhang +2
Low-bit post-training quantization (PTQ) is a pivotal technique for deploying Vision-Language Models (VLMs) on resource-constrained devices. However, existing PTQ methods often deg…
Towards Joint Quantization and Token Pruning of Vision-Language Models
Xinqing Li, Xin He, Xindong Zhang +3
Deploying Vision-Language Models (VLMs) under aggressive low-bit inference remains challenging because inference cost is dominated by the long visual-token prefix during prefill an…
NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models: Datasets, Methods and Results
Xin Li, Jiachao Gong, Xijun Wang +75
This paper presents an overview of the NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models. This challenge utilizes a new short-form UGC (S-…
NTIRE 2024 Restore Any Image Model (RAIM) in the Wild Challenge
Jie Liang, Radu Timofte, Qiaosi Yi +6
In this paper, we review the NTIRE 2024 challenge on Restore Any Image Model (RAIM) in the Wild. The RAIM challenge constructed a benchmark for image restoration in the wild, inclu…
UniVS: Unified and Universal Video Segmentation with Prompts as Queries
Minghan Li, Shuai Li, Xindong Zhang +1
Despite the recent advances in unified image segmentation (IS), developing a unified video segmentation (VS) model remains a challenge. This is mainly because generic category-spec…