activity
20192022
most citedEnd-to-End Multi-Channel Speech Separation

80 citations · 199 across the 17 of their papers we have counts for

collaborators
Showing eess.ASShow all

14 papers · 1 filter

eess.AS202215 cited

Ultra Fast Speech Separation Model with Teacher Student Learning

Sanyuan Chen, Yu Wu, Zhuo Chen +5

Transformer has been successfully applied to speech separation recently with its strong long-dependency modeling capacity using a self-attention mechanism. However, Transformer ten…

eess.AS2021

Continuous Speech Separation with Recurrent Selective Attention Network

Yixuan Zhang, Zhuo Chen, Jian Wu +4

While permutation invariant training (PIT) based continuous speech separation (CSS) significantly improves the conversation transcription accuracy, it often suffers from speech lea…

eess.AS2021

Investigation of Practical Aspects of Single Channel Speech Separation for ASR

Jian Wu, Zhuo Chen, Sanyuan Chen +5

Speech separation has been successfully applied as a frontend processing module of conversation transcription systems thanks to its ability to handle overlapped speech and its flex…

eess.AS20211 cited

A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form Audio

Naoyuki Kanda, Xiong Xiao, Jian Wu +6

Speaker-attributed automatic speech recognition (SA-ASR) is a task to recognize "who spoke what" from multi-talker recordings. An SA-ASR system usually consists of multiple modules…

eess.AS2021

Sequence-level Confidence Classifier for ASR Utterance Accuracy and Application to Acoustic Models

Amber Afshan, Kshitiz Kumar, Jian Wu

Scores from traditional confidence classifiers (CCs) in automatic speech recognition (ASR) systems lack universal interpretation and vary with updates to the underlying confidence…

eess.AS202112 cited

Speaker attribution with voice profiles by graph-based semi-supervised learning

Jixuan Wang, Xiong Xiao, Jian Wu +3

Speaker attribution is required in many real-world applications, such as meeting transcription, where speaker identity is assigned to each utterance according to speaker voice prof…