1 paper · 1 filter
Raphi Kang, Yue Song, Georgia Gkioxari +1
Contrastive Language-Image Pre-Training (CLIP) is a popular method for learning multimodal latent spaces with well-organized semantics. Despite its wide range of applications, CLIP…