13 citations · 18 across the 2 of their papers we have counts for
2 papers
cs.CV2022★ 5 cited
Learning to Answer Questions in Dynamic Audio-Visual Scenarios
Guangyao Li, Yake Wei, Yapeng Tian +3
In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in vid…
cs.CV2022★ 13 cited
Balanced Multimodal Learning via On-the-fly Gradient Modulation
Xiaokang Peng, Yake Wei, Andong Deng +2
Multimodal learning helps to comprehensively understand the world, by integrating different senses. Accordingly, multiple input modalities are expected to boost model performance,…