4 papers
Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
Ao Zhou, Zebo Gu, Tenghao Sun +6
Multimodal Large Language Models (MLLMs) have demonstrated remarkable multimodal understanding capabilities in Visual Question Answering (VQA) tasks by integrating visual and textu…
Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction
Ao Zhou, Mingsheng Tu, Luping Wang +5
Social Media Popularity Prediction is a complex multimodal task that requires effective integration of images, text, and structured information. However, current approaches suffer…
Batch Selection for Multi-Label Classification Guided by Uncertainty and Dynamic Label Correlations
Ao Zhou, Bin Liu, Jin Wang +1
The accuracy of deep neural networks is significantly influenced by the effectiveness of mini-batch construction during training. In single-label scenarios, such as binary and mult…
AEMLO: AutoEncoder-Guided Multi-Label Oversampling
Ao Zhou, Bin Liu, Jin Wang +2
Class imbalance significantly impacts the performance of multi-label classifiers. Oversampling is one of the most popular approaches, as it augments instances associated with less…