1 paper
Zhi Li, Hau Phan, Matthew Emigh +1
Vision-language co-embedding networks, such as CLIP, provide a latent embedding space with semantic information that is useful for downstream tasks. We hypothesize that the embeddi…