1 paper
Jiawei Ma, Po-Yao Huang, Saining Xie +5
The success of contrastive language-image pretraining (CLIP) relies on the supervision from the pairing between images and captions, which tends to be noisy in web-crawled data. We…