1 paper
Emanuele Bugliarello, Laurent Sartran, Aishwarya Agrawal +2
While pretraining on large-scale image-text data from the Web has facilitated rapid progress on many vision-and-language (V&L) tasks, recent work has demonstrated that pretrained m…