1 paper
Hao Zheng, Shunzhi Yang, Zhuoxin He +2
Pre-trained Vision-Language Models (VLMs) such as CLIP have shown excellent generalization abilities. However, adapting these large-scale models to downstream tasks while preservin…