31 citations · 31 across the 1 of their papers we have counts for
1 paper · 1 filter
Zifeng Wang, Zhenbang Wu, Dinesh Agarwal +1
Existing vision-text contrastive learning like CLIP aims to match the paired image and caption embeddings while pushing others apart, which improves representation transferability…