1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.SD2025
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Ming Gao, Shilong Wu, Hang Chen +6
Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted a…
cs.MM2024
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
Mao-Kui He, Jun Du, Shu-Tong Niu +2
In this paper, we propose a quality-aware end-to-end audio-visual neural speaker diarization framework, which comprises three key techniques. First, our audio-visual model takes bo…
eess.AS2023★ 1 cited
The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction
Shilong Wu, Chenxi Wang, Hang Chen +13
Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advan…