collaborators

5 papers

cs.CV2025

GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs

Yuxiang Duan, Ao Li, Yingqin Li +2

Multimodal large language models (MLLMs) have shown remarkable capabilities in a wide range of vision-language tasks. However, the large number of visual tokens introduces signific…

cs.SI2025

From Individuals to Crowds: Dual-Level Public Response Prediction in Social Media

Jinghui Zhang, Kaiyang Wan, Longwei Xu +3

Public response prediction is critical for understanding how individuals or groups might react to specific events, policies, or social phenomena, making it highly valuable for cris…

cs.CV2025

TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model

Ao Li, Yuxiang Duan, Jinghui Zhang +5

Large Vision-Language Models (LVLMs) have advanced multimodal learning but face high computational costs due to the large number of visual tokens, motivating token pruning to impro…

cs.CV2025

Modeling Variants of Prompts for Vision-Language Models

Ao Li, Zongfang Liu, Xinhua Li +3

Large pre-trained vision-language models (VLMs) offer a promising approach to leveraging human language for enhancing downstream tasks. However, VLMs such as CLIP face significant…

cs.CL2024

EmoVerse: Exploring Multimodal Large Language Models for Sentiment and Emotion Understanding

Ao Li, Longwei Xu, Chen Ling +2

Sentiment and emotion understanding are essential to applications such as human-computer interaction and depression detection. While Multimodal Large Language Models (MLLMs) demons…