2 citations · 2 across the 3 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
HAP: Head-Adaptive Visual Token Pruning via Cross-Modal Alignment
Yuanhao Sun, Huawei Ji, Yuan Jin +3
Recent Vision-Language Models encode high-resolution images into long visual token sequences, incurring prohibitive prefill costs. To compress them, existing methods score each vis…
cs.CV2026
ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding
Yuanhao Sun, Huawei Ji, Jiaxin Ding +2
Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising objec…