1 paper
Xiangshuo Qiao, Xianxin Li, Xiaozhe Qu +5
Vision-Language Models pre-trained on large-scale image-text datasets have shown superior performance in downstream tasks such as image retrieval. Most of the images for pre-traini…