823 citations · 859 across the 5 of their papers we have counts for
1 paper · 1 filter
Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare +3
Large-scale vision and language representation learning has shown promising improvements on various vision-language tasks. Most existing methods employ a transformer-based multimod…