2 citations · 2 across the 4 of their papers we have counts for
1 paper · 1 filter
Weiquan Huang, Aoqi Wu, Yifan Yang +10
CLIP is a seminal multimodal model that maps images and text into a shared representation space through contrastive learning on billions of image-caption pairs. Inspired by the rap…