13 citations · 19 across the 4 of their papers we have counts for
4 papers
Learning to Answer Questions in Dynamic Audio-Visual Scenarios
Guangyao Li, Yake Wei, Yapeng Tian +3
In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in vid…
Balanced Multimodal Learning via On-the-fly Gradient Modulation
Xiaokang Peng, Yake Wei, Andong Deng +2
Multimodal learning helps to comprehensively understand the world, by integrating different senses. Accordingly, multiple input modalities are expected to boost model performance,…
SeCo: Separating Unknown Musical Visual Sounds with Consistency Guidance
Xinchi Zhou, Dongzhan Zhou, Wanli Ouyang +3
Recent years have witnessed the success of deep learning on the visual sound separation task. However, existing works follow similar settings where the training and testing dataset…
Visual Sound Localization in the Wild by Cross-Modal Interference Erasing
Xian Liu, Rui Qian, Hang Zhou +5
The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real-world scenarios, audios ar…