1 paper · 1 filter
Kwun Ho Ngan, Saman Sadeghi Afgeh, Joe Townsend +1
Contrastive vision-language models continue to be the dominant approach for image-text retrieval. Contrastive Language-Image Pre-training (CLIP) trains two neural networks to align…