1 paper
Lu Yu, Haiyang Zhang, Changsheng Xu
Due to the impressive zero-shot capabilities, pre-trained vision-language models (e.g., CLIP), have attracted widespread attention and adoption across various domains. Nonetheless,…