4 citations · 7 across the 4 of their papers we have counts for
5 papers
Self-supervised learning with bi-label masked speech prediction for streaming multi-talker speech recognition
Zili Huang, Zhuo Chen, Naoyuki Kanda +6
Self-supervised learning (SSL), which utilizes the input data itself for representation learning, has achieved state-of-the-art results for various downstream speech tasks. However…
Continuous Speech Separation with Recurrent Selective Attention Network
Yixuan Zhang, Zhuo Chen, Jian Wu +4
While permutation invariant training (PIT) based continuous speech separation (CSS) significantly improves the conversation transcription accuracy, it often suffers from speech lea…
Efficient End-to-End Speech Recognition Using Performers in Conformers
Peidong Wang, DeLiang Wang
On-device end-to-end speech recognition poses a high requirement on model efficiency. Most prior works improve the efficiency by reducing model sizes. We propose to reduce the comp…
Speaker Separation Using Speaker Inventories and Estimated Speech
Peidong Wang, Zhuo Chen, DeLiang Wang +2
We propose speaker separation using speaker inventories and estimated speech (SSUSIES), a framework leveraging speaker profiles and estimated speech for speaker separation. SSUSIES…
Bridging the Gap Between Monaural Speech Enhancement and Recognition with Distortion-Independent Acoustic Modeling
Peidong Wang, Ke Tan, DeLiang Wang
Monaural speech enhancement has made dramatic advances since the introduction of deep learning a few years ago. Although enhanced speech has been demonstrated to have better intell…