13 citations · 13 across the 1 of their papers we have counts for
1 paper
Xiaofeng Yang, Fayao Liu, Guosheng Lin
Current vision language pretraining models are dominated by methods using region visual features extracted from object detectors. Given their good performance, the extract-then-pro…