292 citations
- University of WashingtonUS14 papers
- Allen InstituteUS7 papers
- California University of PennsylvaniaUS7 papers
- University of PennsylvaniaUS7 papers
- Bar-Ilan UniversityIL4 papers
- Carnegie Mellon UniversityUS4 papers
- Massachusetts Institute of TechnologyUS3 papers
- Northwestern UniversityUS3 papers
- Cornell UniversityUS2 papers
- Hebrew University of JerusalemIL2 papers
- Indiana UniversityUS2 papers
- Stanford UniversityUS2 papers
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2022
MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound
Rowan Zellers, Jiasen Lu, Ximing Lu +7
As humans, we navigate a multimodal world, building a holistic understanding from all our senses. We introduce MERLOT Reserve, a model that represents videos jointly over time -- t…
cs.CV2020★ 24 cited
X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers
Jaemin Cho, Jiasen Lu, Dustin Schwenk +2
Mirroring the success of masked language models, vision-and-language counterparts like ViLBERT, LXMERT and UNITER have achieved state of the art performance on a variety of multimo…