29 citations · 57 across the 15 of their papers we have counts for
1 paper · 2 filters
Jiaxi Gu, Xiaojun Meng, Guansong Lu +9
Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datase…