1 paper
Hyunjae Kim, Seunghyun Yoon, Trung Bui +4
Contrastive language-image pre-training (CLIP) models have demonstrated considerable success across various vision-language tasks, such as text-to-image retrieval, where the model…