6 papers
Dynamic Distribution-Aware Uncertainty Tracking in Vision-Language Representation Learning
Ao Zhou, Zhiwei Jiang, Zifeng Cheng +4
Uncertainty Quantification (UQ) aims to measure the reliability of model predictions, serving as a critical safeguard for deploying Vision-Language Models (VLMs) in safety-critical…
Multi-Label Test-Time Adaptation with Bayesian Conditional Priors
Qiru Li, Ao Zhou, Zhiwei Jiang +4
Multi-label recognition with frozen Vision-Language Models (VLMs) is brittle under distribution shift: standard zero-shot inference scores labels independently, ignoring co-occurre…
Addressing Imbalance in Multi-Label Data via Label-Specific Distance-based Oversampling
Bin Liu, Jun Wu, Haoyu Peng +4
The complex imbalanced label distribution poses a crucial challenge to multi-label classification, as most classifiers are biased towards the majority class and high-frequent label…
Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
Ao Zhou, Zebo Gu, Tenghao Sun +6
Multimodal Large Language Models (MLLMs) have demonstrated remarkable multimodal understanding capabilities in Visual Question Answering (VQA) tasks by integrating visual and textu…
Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction
Ao Zhou, Mingsheng Tu, Luping Wang +5
Social Media Popularity Prediction is a complex multimodal task that requires effective integration of images, text, and structured information. However, current approaches suffer…
Batch Selection for Multi-Label Classification Guided by Uncertainty and Dynamic Label Correlations
Ao Zhou, Bin Liu, Jin Wang +1
The accuracy of deep neural networks is significantly influenced by the effectiveness of mini-batch construction during training. In single-label scenarios, such as binary and mult…