2 citations · 2 across the 1 of their papers we have counts for
1 paper
Weiquan Huang, Aoqi Wu, Yifan Yang +10
CLIP is a seminal multimodal model that maps images and text into a shared representation space through contrastive learning on billions of image-caption pairs. Inspired by the rap…