2 papers
cs.CV2026
VIG: Visual Information Gain as a Reward Signal for Multimodal Chain-of-Thought Compression
Wen Luo, Xiaohan Yi, Xiaotao Huang +1
Multimodal large reasoning models often rely on long Chain-of-Thought (CoT) traces in which a substantial fraction of tokens, such as repeated visual descriptions, self-reflection,…
cs.CV2026
ViTCoP: Accelerating Large Vision-Language Models via Visual and Textual Semantic Collaborative Pruning
Wen Luo, Peng Chen, Xiaotao Huang +1
Large Vision-Language Models (LVLMs) incur high computational costs due to significant redundancy in their visual tokens. To effectively reduce this cost, researchers have proposed…