3 citations · 3 across the 1 of their papers we have counts for
1 paper
Size Wu, Wenwei Zhang, Sheng Jin +2
Pre-trained vision-language models (VLMs) learn to align vision and language representations on large-scale datasets, where each image-text pair usually contains a bag of semantic…