activity
20202025
most citedUSTC-NELSLIP System Description for DIHARD-III Challenge

20 citations · 23 across the 7 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2023

Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding with Sequence-to-Sequence Architecture

Gaobin Yang, Maokui He, Shutong Niu +6

We propose a novel neural speaker diarization system using memory-aware multi-speaker embedding with sequence-to-sequence architecture (NSD-MS2S), which integrates the strengths of…

eess.AS2023

The USTC-NERCSLIP Systems for the CHiME-7 DASR Challenge

Ruoyu Wang, Maokui He, Jun Du +16

This technical report details our submission system to the CHiME-7 DASR Challenge, which focuses on speaker diarization and speech recognition under complex multi-speaker scenarios…

eess.AS20231 cited

Semi-supervised multi-channel speaker diarization with cross-channel attention

Shilong Wu, Jun Du, Maokui He +4

Most neural speaker diarization systems rely on sufficient manual training data labels, which are hard to collect under real-world scenarios. This paper proposes a semi-supervised…

eess.AS2022

The USTC-Ximalaya system for the ICASSP 2022 multi-channel multi-party meeting transcription (M2MeT) challenge

Maokui He, Xiang Lv, Weilin Zhou +8

We propose two improvements to target-speaker voice activity detection (TS-VAD), the core component in our proposed speaker diarization system that was submitted to the 2022 Multi-…

eess.AS20212 cited

Target-speaker Voice Activity Detection with Improved I-Vector Estimation for Unknown Number of Speaker

Maokui He, Desh Raj, Zili Huang +3

Target-speaker voice activity detection (TS-VAD) has recently shown promising results for speaker diarization on highly overlapped speech. However, the original model requires a fi…

eess.AS2020

Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis

Desh Raj, Pavel Denisov, Zhuo Chen +11

Multi-speaker speech recognition of unsegmented recordings has diverse applications such as meeting transcription and automatic subtitle generation. With technical advances in syst…