3 papers
cs.CV2026
Q Cache: Visual Attention is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model
Jiedong Zhuang, Lu Lu, Ming Dai +4
Multimodal large language models (MLLMs) are plagued by exorbitant inference costs attributable to the profusion of visual tokens within the vision encoder. The redundant visual to…
cs.CL2026
Toward General Semantic Chunking: A Discriminative Framework for Ultra-Long Documents
Kaifeng Wu, Junyan Wu, Qiang Liu +2
Long-document topic segmentation plays an important role in information retrieval and document understanding, yet existing methods still show clear shortcomings in ultra-long text…
cs.CV2024
ST: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming
Jiedong Zhuang, Lu Lu, Ming Dai +4
Multimodal large language models (MLLMs) enhance their perceptual capabilities by integrating visual and textual information. However, processing the massive number of visual token…