110 citations · 244 across the 5 of their papers we have counts for
1 paper · 2 filters
Ming Yan, Haiyang Xu, Chenliang Li +4
Existing approaches to vision-language pre-training (VLP) heavily rely on an object detector based on bounding boxes (regions), where salient objects are first detected from images…