Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models
Jian Chen, Wenye Ma, Penghang Liu +7
Multimodal Large Language Models (MLLMs) have achieved remarkable visual reasoning abilities in natural images, text-rich documents, and graphic designs. However, their ability to…
cs.CV2025
MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion Recognition
Jian Chen, Yuxuan Hu, Haifeng Lu +4
Although pre-trained visual models with text have demonstrated strong capabilities in visual feature extraction, sticker emotion understanding remains challenging due to its relian…