6 citations · 17 across the 6 of their papers we have counts for
1 paper · 1 filter
Soravit Changpinyo, Piyush Sharma, Nan Ding +1
The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. Howev…