1 citations · 1 across the 1 of their papers we have counts for
1 paper
Yanan Sun, Zihan Zhong, Qi Fan +2
Large-scale joint training of multimodal models, e.g., CLIP, have demonstrated great performance in many vision-language tasks. However, image-text pairs for pre-training are restr…