38 citations · 50 across the 3 of their papers we have counts for
1 paper · 1 filter
Siqi Sun, Yen-Chun Chen, Linjie Li +3
Multimodal pre-training has propelled great advancement in vision-and-language research. These large-scale pre-trained models, although successful, fatefully suffer from slow infer…