20 citations · 34 across the 2 of their papers we have counts for
1 paper · 1 filter
Qiaolin Xia, Haoyang Huang, Nan Duan +7
While many BERT-based cross-modal pre-trained models produce excellent results on downstream understanding tasks like image-text retrieval and VQA, they cannot be applied to genera…