4 papers
EvoVLMA: Evolutionary Vision-Language Model Adaptation
Kun Ding, Ying Wang, Shiming Xiang
Pre-trained Vision-Language Models (VLMs) have been exploited in various Computer Vision tasks (e.g., few-shot recognition) via model adaptation, such as prompt tuning and adapters…
Zero-shot Generalizable Incremental Learning for Vision-Language Object Detection
Jieren Deng, Haojian Zhang, Kun Ding +3
This paper presents Incremental Vision-Language Object Detection (IVLOD), a novel learning task designed to incrementally adapt pre-trained Vision-Language Object Detection Models…
A Survey of Low-shot Vision-Language Model Adaptation via Representer Theorem
Kun Ding, Ying Wang, Gaofeng Meng +1
The advent of pre-trained vision-language foundation models has revolutionized the field of zero/few-shot (i.e., low-shot) image recognition. The key challenge to address under the…
Calibrated Cache Model for Few-Shot Vision-Language Model Adaptation
Kun Ding, Qiang Yu, Haojian Zhang +2
Cache-based approaches stand out as both effective and efficient for adapting vision-language models (VLMs). Nonetheless, the existing cache model overlooks three crucial aspects.…