attention dynamics 1efficiency 1multimodal large language models 1token pruning 1training-free methods 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models
Jie Ma, Zhike Qiu, Jie Gao +4
The paper introduces Trend-aware Pruning, a training‑free method that models the temporal dynamics of attention to selectively keep visual tokens that become important in deeper la…
cs.CV2026
Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs
Jie Ma, Zhike Qiu, Jiayi Ji +2
Multimodal Large Language Models (MLLMs) face a significant inference bottleneck due to the quadratic computational cost of self-attention over long visual token sequences. However…
cs.CV2026
Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models
Jie Ma, Yihang Liu, Zhike Qiu +2
Are low-attention visual tokens truly redundant in vision-language reasoning? Existing pruning methods often assume so, ranking visual tokens by shallow text-to-image attention and…