5 citations · 11 across the 4 of their papers we have counts for
4 papers
HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification
Shuyi Ouyang, Hongyi Wang, Ziwei Niu +6
The task of multi-label image classification involves recognizing multiple objects within a single image. Considering both valuable semantic information contained in the labels and…
A Survey on Domain Generalization for Medical Image Analysis
Ziwei Niu, Shuyi Ouyang, Shiao Xie +2
Medical Image Analysis (MedIA) has emerged as a crucial tool in computer-aided diagnosis systems, particularly with the advancement of deep learning (DL) in recent years. However,…
Memory-Inspired Temporal Prompt Interaction for Text-Image Classification
Xinyao Yu, Hao Sun, Ziwei Niu +4
In recent years, large-scale pre-trained multimodal models (LMM) generally emerge to integrate the vision and language modalities, achieving considerable success in various natural…
Modality-invariant and Specific Prompting for Multimodal Human Perception Understanding
Hao Sun, Ziwei Niu, Xinyao Yu +3
Understanding human perceptions presents a formidable multimodal challenge for computers, encompassing aspects such as sentiment tendencies and sense of humor. While various method…