5 papers
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
Yuxiang Duan, Ao Li, Yingqin Li +2
Multimodal large language models (MLLMs) have shown remarkable capabilities in a wide range of vision-language tasks. However, the large number of visual tokens introduces signific…
From Individuals to Crowds: Dual-Level Public Response Prediction in Social Media
Jinghui Zhang, Kaiyang Wan, Longwei Xu +3
Public response prediction is critical for understanding how individuals or groups might react to specific events, policies, or social phenomena, making it highly valuable for cris…
TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
Ao Li, Yuxiang Duan, Jinghui Zhang +5
Large Vision-Language Models (LVLMs) have advanced multimodal learning but face high computational costs due to the large number of visual tokens, motivating token pruning to impro…
Modeling Variants of Prompts for Vision-Language Models
Ao Li, Zongfang Liu, Xinhua Li +3
Large pre-trained vision-language models (VLMs) offer a promising approach to leveraging human language for enhancing downstream tasks. However, VLMs such as CLIP face significant…
EmoVerse: Exploring Multimodal Large Language Models for Sentiment and Emotion Understanding
Ao Li, Longwei Xu, Chen Ling +2
Sentiment and emotion understanding are essential to applications such as human-computer interaction and depression detection. While Multimodal Large Language Models (MLLMs) demons…