activity
20182025
most citedA Real-time Speaker Diarization System Based on Spatial Spectrum

17 citations · 49 across the 12 of their papers we have counts for

collaborators

12 papers

cs.MM20221 cited

MMSpeech: Multi-modal Multi-task Encoder-Decoder Pre-training for Speech Recognition

Xiaohuan Zhou, Jiaming Wang, Zeyu Cui +4

In this paper, we propose a novel multi-modal multi-task encoder-decoder pre-training framework (MMSpeech) for Mandarin automatic speech recognition (ASR), which employs both unlab…

cs.SD20221 cited

Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis

Zhihao Du, Shiliang Zhang, Siqi Zheng +1

Recently, hybrid systems of clustering and neural diarization models have been successfully applied in multi-party meeting analysis. However, current models always treat overlapped…

cs.SD20221 cited

Speaker Embedding-aware Neural Diarization: an Efficient Framework for Overlapping Speech Diarization in Meeting Scenarios

Zhihao Du, Shiliang Zhang, Siqi Zheng +1

Overlapping speech diarization has been traditionally treated as a multi-label classification problem. In this paper, we reformulate this task as a single-label prediction problem…

cs.SD2022

Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge

Fan Yu, Shiliang Zhang, Pengcheng Guo +13

The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologie…

eess.AS2022

ProsoSpeech: Enhancing Prosody With Quantized Vector Pre-training in Text-to-Speech

Yi Ren, Ming Lei, Zhiying Huang +4

Expressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted p…

cs.SD2021

BeamTransformer: Microphone Array-based Overlapping Speech Detection

Siqi Zheng, Shiliang Zhang, Weilong Huang +5

We propose BeamTransformer, an efficient architecture to leverage beamformer's edge in spatial filtering and transformer's capability in context sequence modeling. BeamTransformer…