1 citations · 1 across the 9 of their papers we have counts for
1 paper · 1 filter
Sweta Mahajan, Sukrut Rao, Jiahao Xie +2
Vision-language models such as CLIP are highly useful for diverse tasks due to their shared image-text embedding space. Despite this, the image and text embeddings are often poorly…