6 citations · 10 across the 2 of their papers we have counts for
1 paper · 1 filter
Yuhao Cui, Zhou Yu, Chunqi Wang +4
Vision-and-language pretraining (VLP) aims to learn generic multimodal representations from massive image-text pairs. While various successful attempts have been proposed, learning…