35 citations · 41 across the 7 of their papers we have counts for
6 papers · 1 filter
Stutter-TTS: Controlled Synthesis and Improved Recognition of Stuttered Speech
Xin Zhang, Iván Vallés-Pérez, Andreas Stolcke +5
Stuttering is a speech disorder where the natural flow of speech is interrupted by blocks, repetitions or prolongations of syllables, words and phrases. The majority of existing au…
CoDERT: Distilling Encoder Representations with Co-learning for Transducer-based Speech Recognition
Rupak Vignesh Swaminathan, Brian King, Grant P. Strimel +2
We propose a simple yet effective method to compress an RNN-Transducer (RNN-T) through the well-known knowledge distillation paradigm. We show that the transducer's encoder outputs…
Do as I mean, not as I say: Sequence Loss Training for Spoken Language Understanding
Milind Rao, Pranav Dheram, Gautam Tiwari +4
Spoken language understanding (SLU) systems extract transcriptions, as well as semantics of intent or named entities from speech, and are essential components of voice activated sy…
Acoustic-To-Word Model Without OOV
Jinyu Li, Guoli Ye, Rui Zhao +2
Recently, the acoustic-to-word model based on the Connectionist Temporal Classification (CTC) criterion was shown as a natural end-to-end model directly targeting words as output u…
Progressive Joint Modeling in Unsupervised Single-channel Overlapped Speech Recognition
Zhehuai Chen, Jasha Droppo, Jinyu Li +1
Unsupervised single-channel overlapped speech recognition is one of the hardest problems in automatic speech recognition (ASR). Permutation invariant training (PIT) is a state of t…
On Training Bi-directional Neural Network Language Model with Noise Contrastive Estimation
Tianxing He, Yu Zhang, Jasha Droppo +1
We propose to train bi-directional neural network language model(NNLM) with noise contrastive estimation(NCE). Experiments are conducted on a rescore task on the PTB data set. It i…