5 papers
MAC: A Benchmark for Multiple Attributes Compositional Zero-Shot Learning
Shuo Xu, Sai Wang, Xinyue Hu +3
Compositional Zero-Shot Learning (CZSL) aims to learn semantic primitives (attributes and objects) from seen compositions and recognize unseen attribute-object compositions. Existi…
Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis
Ruilang Wang, Shuotong Xu, Bowen Liu +3
The scarcity of annotated data in specialized domains such as medical imaging presents significant challenges to training robust vision models. While self-supervised masked image m…
Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt Tuning
LeiLei Ma, Shuo Xu, MingKun Xie +3
Modeling label correlations has always played a pivotal role in multi-label image classification (MLC), attracting significant attention from researchers. However, recent studies h…
Bridging Vision and Language: Optimal Transport-Driven Radiology Report Generation via LLMs
Haifeng Zhao, Yufei Zhang, Leilei Ma +2
Radiology report generation represents a significant application within medical AI, and has achieved impressive results. Concurrently, large language models (LLMs) have demonstrate…
TACOcc:Target-Adaptive Cross-Modal Fusion with Volume Rendering for 3D Semantic Occupancy
Luyao Lei, Shuo Xu, Yifan Bai +1
The performance of multi-modal 3D occupancy prediction is limited by ineffective fusion, mainly due to geometry-semantics mismatch from fixed fusion strategies and surface detail l…