activity
20172022
most citedOn Mean Absolute Error for Deep Neural Network Based Vector-to-Vector Regression

303 citations · 631 across the 26 of their papers we have counts for

collaborators
Showing eess.ASShow all

17 papers · 1 filter

eess.AS20221 cited

Deep Learning Based Audio-Visual Multi-Speaker DOA Estimation Using Permutation-Free Loss Function

Qing Wang, Hang Chen, Ya Jiang +4

In this paper, we propose a deep learning based multi-speaker direction of arrival (DOA) estimation with audio and visual signals by using permutation-free loss function. We first…

eess.AS2022

The USTC-Ximalaya system for the ICASSP 2022 multi-channel multi-party meeting transcription (M2MeT) challenge

Maokui He, Xiang Lv, Weilin Zhou +8

We propose two improvements to target-speaker voice activity detection (TS-VAD), the core component in our proposed speaker diarization system that was submitted to the 2022 Multi-…

eess.AS20212 cited

Target-speaker Voice Activity Detection with Improved I-Vector Estimation for Unknown Number of Speaker

Maokui He, Desh Raj, Zili Huang +3

Target-speaker voice activity detection (TS-VAD) has recently shown promising results for speaker diarization on highly overlapped speech. However, the original model requires a fi…

eess.AS20215 cited

Separation Guided Speaker Diarization in Realistic Mismatched Conditions

Shu-Tong Niu, Jun Du, Lei Sun +1

We propose a separation guided speaker diarization (SGSD) approach by fully utilizing a complementarity of speech separation and speaker clustering. Since the conventional clusteri…

eess.AS2020

The Third DIHARD Diarization Challenge

Neville Ryant, Prachi Singh, Venkat Krishnamohan +6

DIHARD III was the third in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variability in recording equipment, noise condit…

eess.AS2020

Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis

Desh Raj, Pavel Denisov, Zhuo Chen +11

Multi-speaker speech recognition of unsegmented recordings has diverse applications such as meeting transcription and automatic subtitle generation. With technical advances in syst…