17 citations · 55 across the 26 of their papers we have counts for
5 papers · 1 filter
Serialized Output Training for End-to-End Overlapped Speech Recognition
Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang +2
This paper proposes serialized output training (SOT), a novel framework for multi-speaker overlapped speech recognition based on an attention-based encoder-decoder approach. Instea…
A practical two-stage training strategy for multi-stream end-to-end speech recognition
Ruizhi Li, Gregory Sell, Xiaofei Wang +2
The multi-stream paradigm of audio processing, in which several sources are simultaneously considered, has been an active research area for information fusion. Our previous study o…
Exploring Methods for the Automatic Detection of Errors in Manual Transcription
Xiaofei Wang, Jinyi Yang, Ruizhi Li +2
Quality of data plays an important role in most deep learning tasks. In the speech community, transcription of speech recording is indispensable. Since the transcription is usually…
Multi-encoder multi-resolution framework for end-to-end speech recognition
Ruizhi Li, Xiaofei Wang, Sri Harish Mallidi +3
Attention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end Automatic Speech Recognition (ASR). The joint…
Stream attention-based multi-array end-to-end speech recognition
Xiaofei Wang, Ruizhi Li, Sri Harish Mallid +3
Automatic Speech Recognition (ASR) using multiple microphone arrays has achieved great success in the far-field robustness. Taking advantage of all the information that each array…