19 citations · 27 across the 4 of their papers we have counts for
4 papers · 1 filter
Multi-encoder multi-resolution framework for end-to-end speech recognition
Ruizhi Li, Xiaofei Wang, Sri Harish Mallidi +3
Attention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end Automatic Speech Recognition (ASR). The joint…
Stream attention-based multi-array end-to-end speech recognition
Xiaofei Wang, Ruizhi Li, Sri Harish Mallid +3
Automatic Speech Recognition (ASR) using multiple microphone arrays has achieved great success in the far-field robustness. Taking advantage of all the information that each array…
Multilingual sequence-to-sequence speech recognition: architecture, transfer learning, and language modeling
Jaejin Cho, Murali Karthick Baskar, Ruizhi Li +6
Sequence-to-sequence (seq2seq) approach for low-resource ASR is a relatively new direction in speech research. The approach benefits by performing model training without using lexi…
Device-directed Utterance Detection
Sri Harish Mallidi, Roland Maas, Kyle Goehner +3
In this work, we propose a classifier for distinguishing device-directed queries from background speech in the context of interactions with voice assistants. Applications include r…