3 papers
eess.AS2023
Fusion of Audio and Visual Embeddings for Sound Event Localization and Detection
Davide Berghi, Peipei Wu, Jinzheng Zhao +2
Sound event localization and detection (SELD) combines two subtasks: sound event detection (SED) and direction of arrival (DOA) estimation. SELD is usually tackled as an audio-only…
cs.CV2023
CM-PIE: Cross-modal perception for interactive-enhanced audio-visual video parsing
Yaru Chen, Ruohao Guo, Xubo Liu +4
Audio-visual video parsing is the task of categorizing a video at the segment level with weak labels, and predicting them as audible or visible events. Recent methods for this task…
cs.MM2023
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
Jinzheng Zhao, Yong Xu, Xinyuan Qian +6
Audio-visual speaker tracking has drawn increasing attention over the past few years due to its academic values and wide applications. Audio and visual modalities can provide compl…