2 papers
cs.CV2026
Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs
Jouwon Song, Woohyeong Kim, Kyeongbo Kong
Recent high-resolution Multimodal Large Language Models (MLLMs) generate thousands of visual tokens per input, leading to a visual token explosion that introduces severe latency bo…
cs.CV2026
AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models
Changwoo Baek, Jouwon Song, Sohyeon Kim +1
Large Vision-Language Models (LVLMs) have adopted visual token pruning strategies to mitigate substantial computational overhead incurred by extensive visual token sequences. While…