22 citations · 22 across the 2 of their papers we have counts for
2 papers
cs.CV2022
MUGEN: A Playground for Video-Audio-Text Multimodal Understanding and GENeration
Thomas Hayes, Songyang Zhang, Xi Yin +6
Multimodal video-audio-text understanding and generation can benefit from datasets that are narrow but rich. The narrowness allows bite-sized challenges that the research community…
cs.CV2021★ 22 cited
Human-Adversarial Visual Question Answering
Sasha Sheng, Amanpreet Singh, Vedanuj Goswami +4
Performance on the most commonly used Visual Question Answering dataset (VQA v2) is starting to approach human accuracy. However, in interacting with state-of-the-art VQA models, i…