2 papers
eess.AS2025
M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset
Shilong Wu
In the field of speaker diarization, the development of technology is constrained by two problems: insufficient data resources and poor generalization ability of deep learning mode…
cs.SD2025
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Ming Gao, Shilong Wu, Hang Chen +6
Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted a…