13 citations · 19 across the 5 of their papers we have counts for
5 papers · 1 filter
Progressive Spatio-temporal Perception for Audio-Visual Question Answering
Guangyao Li, Wenxuan Hou, Di Hu
Audio-Visual Question Answering (AVQA) task aims to answer questions about different visual objects, sounds, and their associations in videos. Such naturally multi-modal videos are…
Revisiting Pre-training in Audio-Visual Learning
Ruoxuan Feng, Wenke Xia, Di Hu
Pre-training technique has gained tremendous success in enhancing model performance on various tasks, but found to perform worse than training from scratch in some uni-modal situat…
Learning to Answer Questions in Dynamic Audio-Visual Scenarios
Guangyao Li, Yake Wei, Yapeng Tian +3
In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in vid…
Balanced Multimodal Learning via On-the-fly Gradient Modulation
Xiaokang Peng, Yake Wei, Andong Deng +2
Multimodal learning helps to comprehensively understand the world, by integrating different senses. Accordingly, multiple input modalities are expected to boost model performance,…
Visual Sound Localization in the Wild by Cross-Modal Interference Erasing
Xian Liu, Rui Qian, Hang Zhou +5
The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real-world scenarios, audios ar…