7 papers · 1 filter
MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression
Guangheng Yang, Zhenliang Ni, Zhenkai Wu +4
Recently, multimodal large-scale reasoning models have demonstrated remarkable capabilities in solving complex tasks through long Chains-of-Thought (M-CoT). However, excessively lo…
CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering
Mouxiao Huang, Qiangyu Yan, Borui Jiang +1
Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (e.g., CIDEr and SPICE) and L…
SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models
Yaozhi Wen, Jialong Guo, Zhenliang Ni +2
While Vision-Language Models (VLMs) have demonstrated remarkable performance in processing and understanding both text and images, their large parameter sizes lead to significant c…
TinySAM 2: Extreme Memory Compression for Efficient Track Anything Model
Zhaoyuan Ding, Yijing Yang, Han Shu +1
Segment Anything Model 2 (SAM 2) serves as a core foundation model in the field of video segmentation. Building upon the original SAM model, it introduces a memory bank mechanism a…
SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation
Jialiang Kang, Han Shu, Wenshuo Li +2
Speculative Jacobi Decoding (SJD) offers a draft-model-free approach to accelerate autoregressive text-to-image synthesis. However, the high-entropy nature of visual generation yie…
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
Zhenkai Wu, Xiaowen Ma, Zhenliang Ni +4
Vision-language models (VLMs) excel at image understanding tasks, but the large number of visual tokens imposes significant computational costs, hindering deployment on mobile devi…