1 paper
Zhouyao Xie, Nikhil Yadala, Xinyi Chen +1
CLIP (Contrastive Language-Image Pre-Training) is a multimodal neural network trained on (text, image) pairs to predict the most relevant text caption given an image. It has been u…