4 papers
Integrated Structural Prompt Learning for Vision-Language Models
Jiahui Wang, Qin Xu, Bo Jiang +1
Prompt learning methods have significantly extended the transferability of pre-trained Vision-Language Models (VLMs) like CLIP for various downstream tasks. These methods adopt han…
Dynamic Rank Adaptation for Vision-Language Models
Jiahui Wang, Qin Xu, Bo Jiang +1
Pre-trained large vision-language models (VLMs) like CLIP demonstrate impressive generalization ability. Existing prompt-based and adapter-based works have made significant progres…
Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks
Yuhe Ding, Bo Jiang, Aihua Zheng +2
Vision language models (VLMs) like CLIP show stellar zero-shot capability on classification benchmarks. However, selecting the VLM with the highest performance on the unlabeled dow…
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
Yumiao Zhao, Bo Jiang, Xiao Wang +2
Adapter-based tuning methods have shown significant potential in transferring knowledge from pre-trained Vision-Language Models to the downstream tasks. However, after reviewing ex…