2 papers
cs.CL2025
Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and Fusion
Daiqing Wu, Dongbao Yang, Yu Zhou +1
As posts on social media increase rapidly, analyzing the sentiments embedded in image-text pairs has become a popular research topic in recent years. Although existing works achiev…
cs.CV2025
CROP: Contextual Region-Oriented Visual Token Pruning
Jiawei Guo, Feifei Zhai, Pu Jian +2
Current VLM-based VQA methods often process entire images, leading to excessive visual tokens that include redundant information irrelevant to the posed question. This abundance of…