2 citations · 3 across the 8 of their papers we have counts for
5 papers · 1 filter
A practical two-stage training strategy for multi-stream end-to-end speech recognition
Ruizhi Li, Gregory Sell, Xiaofei Wang +2
The multi-stream paradigm of audio processing, in which several sources are simultaneously considered, has been an active research area for information fusion. Our previous study o…
Performance Monitoring for End-to-End Speech Recognition
Ruizhi Li, Gregory Sell, Hynek Hermansky
Measuring performance of an automatic speech recognition (ASR) system without ground-truth could be beneficial in many scenarios, especially with data from unseen domains, where pe…
Exploring Methods for the Automatic Detection of Errors in Manual Transcription
Xiaofei Wang, Jinyi Yang, Ruizhi Li +2
Quality of data plays an important role in most deep learning tasks. In the speech community, transcription of speech recording is indispensable. Since the transcription is usually…
Multi-encoder multi-resolution framework for end-to-end speech recognition
Ruizhi Li, Xiaofei Wang, Sri Harish Mallidi +3
Attention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end Automatic Speech Recognition (ASR). The joint…
Stream attention-based multi-array end-to-end speech recognition
Xiaofei Wang, Ruizhi Li, Sri Harish Mallid +3
Automatic Speech Recognition (ASR) using multiple microphone arrays has achieved great success in the far-field robustness. Taking advantage of all the information that each array…