8 citations · 8 across the 1 of their papers we have counts for
1 paper
Haiyang Xu, Ming Yan, Chenliang Li +4
Vision-language pre-training (VLP) on large-scale image-text pairs has achieved huge success for the cross-modal downstream tasks. The most existing pre-training methods mainly ado…