Showing cs.IRShow all
2 papers · 1 filter
cs.IR2025
Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
Ao Zhou, Zebo Gu, Tenghao Sun +6
Multimodal Large Language Models (MLLMs) have demonstrated remarkable multimodal understanding capabilities in Visual Question Answering (VQA) tasks by integrating visual and textu…
cs.IR2025
Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction
Ao Zhou, Mingsheng Tu, Luping Wang +5
Social Media Popularity Prediction is a complex multimodal task that requires effective integration of images, text, and structured information. However, current approaches suffer…