2 citations · 2 across the 4 of their papers we have counts for
4 papers
Contextual-Utterance Training for Automatic Speech Recognition
Alejandro Gomez-Alanis, Lukas Drude, Andreas Schwarz +2
Recent studies of streaming automatic speech recognition (ASR) recurrent neural network transducer (RNN-T)-based systems have fed the encoder with past contextual information in or…
ConvRNN-T: Convolutional Augmented Recurrent Neural Network Transducers for Streaming Speech Recognition
Martin Radfar, Rohit Barnwal, Rupak Vignesh Swaminathan +4
The recurrent neural network transducer (RNN-T) is a prominent streaming end-to-end (E2E) ASR technology. In RNN-T, the acoustic encoder commonly consists of stacks of LSTMs. Very…
CoDERT: Distilling Encoder Representations with Co-learning for Transducer-based Speech Recognition
Rupak Vignesh Swaminathan, Brian King, Grant P. Strimel +2
We propose a simple yet effective method to compress an RNN-Transducer (RNN-T) through the well-known knowledge distillation paradigm. We show that the transducer's encoder outputs…
Exploiting Large-scale Teacher-Student Training for On-device Acoustic Models
Jing Liu, Rupak Vignesh Swaminathan, Sree Hari Krishnan Parthasarathi +3
We present results from Alexa speech teams on semi-supervised learning (SSL) of acoustic models (AM) with experiments spanning over 3000 hours of GPU time, making our study one of…