4 papers
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
Qi Li, Yanzhe Zhao, Yongxin Zhou +4
Multimodal Large Language Models (MLLMs) have shown immense promise in universal multimodal retrieval, which aims to find relevant items of various modalities for a given query. Ho…
A Novel Graph-Sequence Learning Model for Inductive Text Classification
Zuo Wang, Ye Yuan
Text classification plays an important role in various downstream text-related tasks, such as sentiment analysis, fake news detection, and public opinion analysis. Recently, text c…
Jensen-Shannon Divergence Message-Passing for Rich-Text Graph Representation Learning
Zuo Wang, Ye Yuan
In this paper, we investigate how the widely existing contextual and structural divergence may influence the representation learning in rich-text graphs. To this end, we propose Je…
Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
Wei Tang, Zuo-Zheng Wang, Kun Zhang +2
Long-tailed multi-label visual recognition poses a significant challenge, as images typically contain multiple labels with highly imbalanced class distributions, leading to biased…