99 citations · 144 across the 8 of their papers we have counts for
1 paper · 1 filter
Ming Yan, Haiyang Xu, Chenliang Li +4
Existing approaches to vision-language pre-training (VLP) heavily rely on an object detector based on bounding boxes (regions), where salient objects are first detected from images…