29 citations · 29 across the 1 of their papers we have counts for
1 paper
Jiaxi Gu, Xiaojun Meng, Guansong Lu +9
Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datase…