1 paper
Zhi Chen, Xin Yu, Xiaohui Tao +2
Vision-language models (VLMs) such as CLIP achieve zero-shot transfer across various tasks by pre-training on numerous image-text pairs. These models often benefit from using an en…