1 citations · 1 across the 1 of their papers we have counts for
1 paper
Masoud Monajatipoor, Mozhdeh Rouhsedaghat, Liunian Harold Li +4
Vision-and-language(V&L) models take image and text as input and learn to capture the associations between them. Prior studies show that pre-trained V&L models can significantly im…