9 citations · 36 across the 13 of their papers we have counts for
6 papers · 1 filter
REDAT: Accent-Invariant Representation for End-to-End ASR by Domain Adversarial Training with Relabeling
Hu Hu, Xuesong Yang, Zeynab Raeesy +6
Accents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DA…
Efficient minimum word error rate training of RNN-Transducer for end-to-end speech recognition
Jinxi Guo, Gautam Tiwari, Jasha Droppo +4
In this work, we propose a novel and efficient minimum word error rate (MWER) training method for RNN-Transducer (RNN-T). Unlike previous work on this topic, which performs on-the-…
Streaming ResLSTM with Causal Mean Aggregation for Device-Directed Utterance Detection
Xiaosu Tong, Che-Wei Huang, Sri Harish Mallidi +5
In this paper, we propose a streaming model to distinguish voice queries intended for a smart-home device from background speech. The proposed model consists of multiple CNN layers…
Streaming End-to-End Bilingual ASR Systems with Joint Language Identification
Surabhi Punjabi, Harish Arsikere, Zeynab Raeesy +11
Multilingual ASR technology simplifies model training and deployment, but its accuracy is known to depend on the availability of language information at runtime. Since language ide…
Multi-view Frequency LSTM: An Efficient Frontend for Automatic Speech Recognition
Maarten Van Segbroeck, Harish Mallidih, Brian King +3
Acoustic models in real-time speech recognition systems typically stack multiple unidirectional LSTM layers to process the acoustic frames over time. Performance improvements over…
Streaming Language Identification using Combination of Acoustic Representations and ASR Hypotheses
Chander Chandak, Zeynab Raeesy, Ariya Rastrow +5
This paper presents our modeling and architecture approaches for building a highly accurate low-latency language identification system to support multilingual spoken queries for vo…