303 citations · 631 across the 26 of their papers we have counts for
17 papers · 1 filter
Deep Learning Based Audio-Visual Multi-Speaker DOA Estimation Using Permutation-Free Loss Function
Qing Wang, Hang Chen, Ya Jiang +4
In this paper, we propose a deep learning based multi-speaker direction of arrival (DOA) estimation with audio and visual signals by using permutation-free loss function. We first…
The USTC-Ximalaya system for the ICASSP 2022 multi-channel multi-party meeting transcription (M2MeT) challenge
Maokui He, Xiang Lv, Weilin Zhou +8
We propose two improvements to target-speaker voice activity detection (TS-VAD), the core component in our proposed speaker diarization system that was submitted to the 2022 Multi-…
Target-speaker Voice Activity Detection with Improved I-Vector Estimation for Unknown Number of Speaker
Maokui He, Desh Raj, Zili Huang +3
Target-speaker voice activity detection (TS-VAD) has recently shown promising results for speaker diarization on highly overlapped speech. However, the original model requires a fi…
Separation Guided Speaker Diarization in Realistic Mismatched Conditions
Shu-Tong Niu, Jun Du, Lei Sun +1
We propose a separation guided speaker diarization (SGSD) approach by fully utilizing a complementarity of speech separation and speaker clustering. Since the conventional clusteri…
The Third DIHARD Diarization Challenge
Neville Ryant, Prachi Singh, Venkat Krishnamohan +6
DIHARD III was the third in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variability in recording equipment, noise condit…
Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis
Desh Raj, Pavel Denisov, Zhuo Chen +11
Multi-speaker speech recognition of unsegmented recordings has diverse applications such as meeting transcription and automatic subtitle generation. With technical advances in syst…