99 citations · 144 across the 8 of their papers we have counts for
1 paper · 1 filter
Haiyang Xu, Ming Yan, Chenliang Li +4
Vision-language pre-training (VLP) on large-scale image-text pairs has achieved huge success for the cross-modal downstream tasks. The most existing pre-training methods mainly ado…