activity
20162024
most citedStreaming End-to-End Bilingual ASR Systems with Joint Language Identification

9 citations · 36 across the 13 of their papers we have counts for

collaborators
Showing 2020Show all

6 papers · 1 filter

eess.AS2020★ 2 cited

REDAT: Accent-Invariant Representation for End-to-End ASR by Domain Adversarial Training with Relabeling

Hu Hu, Xuesong Yang, Zeynab Raeesy +6

Accents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DA…

eess.AS2020★ 1 cited

Efficient minimum word error rate training of RNN-Transducer for end-to-end speech recognition

Jinxi Guo, Gautam Tiwari, Jasha Droppo +4

In this work, we propose a novel and efficient minimum word error rate (MWER) training method for RNN-Transducer (RNN-T). Unlike previous work on this topic, which performs on-the-…

eess.AS2020★ 4 cited

Streaming ResLSTM with Causal Mean Aggregation for Device-Directed Utterance Detection

Xiaosu Tong, Che-Wei Huang, Sri Harish Mallidi +5

In this paper, we propose a streaming model to distinguish voice queries intended for a smart-home device from background speech. The proposed model consists of multiple CNN layers…

eess.AS2020★ 9 cited

Streaming End-to-End Bilingual ASR Systems with Joint Language Identification

Surabhi Punjabi, Harish Arsikere, Zeynab Raeesy +11

Multilingual ASR technology simplifies model training and deployment, but its accuracy is known to depend on the availability of language information at runtime. Since language ide…

eess.AS2020★ 3 cited

Multi-view Frequency LSTM: An Efficient Frontend for Automatic Speech Recognition

Maarten Van Segbroeck, Harish Mallidih, Brian King +3

Acoustic models in real-time speech recognition systems typically stack multiple unidirectional LSTM layers to process the acoustic frames over time. Performance improvements over…

eess.AS2020★ 9 cited

Streaming Language Identification using Combination of Acoustic Representations and ASR Hypotheses

Chander Chandak, Zeynab Raeesy, Ariya Rastrow +5

This paper presents our modeling and architecture approaches for building a highly accurate low-latency language identification system to support multilingual spoken queries for vo…