8 citations · 10 across the 5 of their papers we have counts for
1 paper · 1 filter
Manling Li, Ruochen Xu, Shuohang Wang +6
Vision-language (V+L) pretraining models have achieved great success in supporting multimedia applications by understanding the alignments between images and text. While existing v…