4 papers
CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction
Jiadong Wang, Ke Zhang, Xinyuan Qian +3
Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Alt…
Region-Specific Audio Tagging for Spatial Sound
Jinzheng Zhao, Yong Xu, Haohe Liu +6
Audio tagging aims to label sound events appearing in an audio recording. In this paper, we propose region-specific audio tagging, a new task which labels sound events in a given r…
Audio-Visual Class-Incremental Learning for Fish Feeding intensity Assessment in Aquaculture
Meng Cui, Xianghu Yue, Xinyuan Qian +5
Fish Feeding Intensity Assessment (FFIA) is crucial in industrial aquaculture management. Recent multi-modal approaches have shown promise in improving FFIA robustness and efficien…
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
Jinzheng Zhao, Yong Xu, Xinyuan Qian +6
Audio-visual speaker tracking has drawn increasing attention over the past few years due to its academic values and wide applications. Audio and visual modalities can provide compl…