315 citations · 955 across the 39 of their papers we have counts for
Showing 2021 · cs.CVShow all
2 papers · 2 filters
cs.CV2021★ 22 cited
Human-Adversarial Visual Question Answering
Sasha Sheng, Amanpreet Singh, Vedanuj Goswami +4
Performance on the most commonly used Visual Question Answering dataset (VQA v2) is starting to approach human accuracy. However, in interacting with state-of-the-art VQA models, i…
cs.CV2021★ 8 cited
VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs
Xudong Lin, Gedas Bertasius, Jue Wang +3
We present \textsc{Vx2Text}, a framework for text generation from multimodal inputs consisting of video plus text, speech, or audio. In order to leverage transformer networks, whic…