4 citations · 4 across the 5 of their papers we have counts for
1 paper · 1 filter
Yuan Yao, Qianyu Chen, Ao Zhang +4
Vision-language pre-training (VLP) has shown impressive performance on a wide range of cross-modal tasks, where VLP models without reliance on object detectors are becoming the mai…