3 papers
cs.CV2025
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
Yuxiang Duan, Ao Li, Yingqin Li +2
Multimodal large language models (MLLMs) have shown remarkable capabilities in a wide range of vision-language tasks. However, the large number of visual tokens introduces signific…
cs.CL2025
EmoVerse: Exploring Multimodal Large Language Models for Sentiment and Emotion Understanding
Ao Li, Longwei Xu, Chen Ling +2
Sentiment and emotion understanding are essential to applications such as human-computer interaction and depression detection. While Multimodal Large Language Models (MLLMs) demons…
cs.CV2025
Modeling Variants of Prompts for Vision-Language Models
Ao Li, Zongfang Liu, Xinhua Li +3
Large pre-trained vision-language models (VLMs) offer a promising approach to leveraging human language for enhancing downstream tasks. However, VLMs such as CLIP face significant…