1 paper
Zheng Zeng, Deepak Sridhar, Nuno Vasconcelos
Vision-language models (VLMs) such as CLIP enable zero-shot classification by comparing image features with text prompts in a shared embedding space. A fundamental property underly…