1 paper
Samuel Lavoie, Polina Kirichenko, Mark Ibrahim +4
There are a thousand ways to caption an image. Contrastive Language Pretraining (CLIP) on the other hand, works by mapping an image and its caption to a single vector -- limiting h…