244 citations · 372 across the 32 of their papers we have counts for
1 paper · 1 filter
Zifeng Wang, Zhenbang Wu, Dinesh Agarwal +1
Existing vision-text contrastive learning like CLIP aims to match the paired image and caption embeddings while pushing others apart, which improves representation transferability…