4 papers · 1 filter
MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs
Zhongyang Li, Yaqian Li, Faming Fang +6
Multimodal large language models (MLLMs) typically employ resampling-based projectors to transform dense visual features into a compact token sequence for language modeling. Most e…
SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs
Zi-Hao Bo, Yaqian Li, Anzhou Hou +6
Mixture-of-Experts (MoE) has become a prevalent backbone for large vision-language models (VLMs), yet how modality-specific signals should guide expert routing remains under-explor…
LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models
Rinyoichi Takezoe, Yaqian Li, Zihao Bo +3
Vision-Language Models (VLMs) have recently demonstrated remarkable capabilities in visual understanding and reasoning, but they also impose significant computational burdens due t…
QMoP: Query Guided Mixture-of-Projector for Efficient Visual Token Compression
Zhongyang Li, Yaqian Li, Faming Fang +6
Multimodal large language models suffer from severe computational and memory bottlenecks, as the number of visual tokens far exceeds that of textual tokens. While recent methods em…