31 citations · 61 across the 7 of their papers we have counts for
1 paper · 1 filter
Zifeng Wang, Zhenbang Wu, Dinesh Agarwal +1
Existing vision-text contrastive learning like CLIP aims to match the paired image and caption embeddings while pushing others apart, which improves representation transferability…