2 papers
cs.AI2026
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin
Yuyao Sun, Tao Deng, Shuang Li +3
Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual…
cs.CV2025
InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
Dongchen Lu, Yuyao Sun, Zilu Zhang +4
Most multimodal large language models (MLLMs) treat visual tokens as "a sequence of text", integrating them with text tokens into a large language model (LLM). However, a great qua…