54 citations · 68 across the 2 of their papers we have counts for
2 papers
cs.CV2021★ 54 cited
MERLOT: Multimodal Neural Script Knowledge Models
Rowan Zellers, Ximing Lu, Jack Hessel +5
As humans, we understand events in the visual world contextually, performing multimodal reasoning across time to make inferences about the past, present, and future. We introduce M…
cs.CV2020★ 14 cited
Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models
Jize Cao, Zhe Gan, Yu Cheng +3
Recent Transformer-based large-scale pre-trained models have revolutionized vision-and-language (V+L) research. Models such as ViLBERT, LXMERT and UNITER have significantly lifted…