5 papers
TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
Ao Li, Yuxiang Duan, Jinghui Zhang +5
Large Vision-Language Models (LVLMs) have advanced multimodal learning but face high computational costs due to the large number of visual tokens, motivating token pruning to impro…
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
Yuxiang Duan, Ao Li, Yingqin Li +2
Multimodal large language models (MLLMs) have shown remarkable capabilities in a wide range of vision-language tasks. However, the large number of visual tokens introduces signific…
From Individuals to Crowds: Dual-Level Public Response Prediction in Social Media
Jinghui Zhang, Kaiyang Wan, Longwei Xu +3
Public response prediction is critical for understanding how individuals or groups might react to specific events, policies, or social phenomena, making it highly valuable for cris…
EmoVerse: Exploring Multimodal Large Language Models for Sentiment and Emotion Understanding
Ao Li, Longwei Xu, Chen Ling +2
Sentiment and emotion understanding are essential to applications such as human-computer interaction and depression detection. While Multimodal Large Language Models (MLLMs) demons…
Modeling Variants of Prompts for Vision-Language Models
Ao Li, Zongfang Liu, Xinhua Li +3
Large pre-trained vision-language models (VLMs) offer a promising approach to leveraging human language for enhancing downstream tasks. However, VLMs such as CLIP face significant…