2 citations · 3 across the 11 of their papers we have counts for
1 paper · 2 filters
Sweta Mahajan, Sukrut Rao, Jiahao Xie +2
Vision-language models such as CLIP are highly useful for diverse tasks due to their shared image-text embedding space. Despite this, the image and text embeddings are often poorly…