6 papers
One Stone, Three Birds: Self-adaptive Optimal Transport for Multi-VLM Selection, Adaptation, and Ensembling
Qiyu Xu, Zhanxuan Hu, Yu Duan +4
Vision-language models (VLMs) enable visual recognition from semantic class descriptions, which makes them attractive when target annotations are scarce or unavailable. Most deploy…
Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation Models
Shuwen Yu, Zhanxuan Hu, Yi Zhao +2
Foundation models have driven rapid progress in computer vision, yet the two dominant paradigms, vision-language foundation models (VLMs) and vision-only foundation models (VFMs),…
[CLS] is Not Enough: Multi-Label Recognition via Patch-Level Inference and Adaptive Aggregation
Akang Wang, Xili Deng, Zhanxuan Hu +3
Vision-Language Models such as CLIP exhibit strong zero-shot recognition capability by aligning images with textual concepts, yet they often underperform on multi-label recognition…
ConInfer: Context-Aware Inference for Training-Free Open-Vocabulary Remote Sensing Segmentation
Wenyang Chen, Zhanxuan Hu, Yaping Zhang +2
Training-free open-vocabulary remote sensing segmentation (OVRSS), empowered by vision-language models, has emerged as a promising paradigm for achieving category-agnostic semantic…
SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation Models
Zhanxuan Hu, Qiyu Xu, Yu Duan +2
Foundation models have attracted widespread attention across domains due to their powerful zero-shot classification capabilities. This work is motivated by two key observations: (1…
A Hidden Stumbling Block in Generalized Category Discovery: Distracted Attention
Qiyu Xu, Zhanxuan Hu, Yu Duan +2
Generalized Category Discovery (GCD) aims to classify unlabeled data from both known and unknown categories by leveraging knowledge from labeled known categories. While existing me…