1 paper
Zhihe Lu, Jiawang Bai, Xin Li +2
Fine-tuning pre-trained vision-language models (VLMs), e.g., CLIP, for the open-world generalization has gained increasing popularity due to its practical value. However, performan…