RNN-T For Latency Controlled ASR With Improved Beam Search
arXiv:1911.01629
Abstract
Neural transducer-based systems such as RNN Transducers (RNN-T) for automatic speech recognition (ASR) blend the individual components of a traditional hybrid ASR systems (acoustic model, language model, punctuation model, inverse text normalization) into one single model. This greatly simplifies training and inference and hence makes RNN-T a desirable choice for ASR systems. In this work, we investigate use of RNN-T in applications that require a tune-able latency budget during inference time. We also improved the decoding speed of the originally proposed RNN-T beam search algorithm. We evaluated our proposed system on English videos ASR dataset and show that neural RNN-T models can achieve comparable WER and better computational efficiency compared to a well tuned hybrid ASR baseline.
References in corpus (1)
Cited by in corpus (6)
- On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition
- Attention-based Transducer for Online Speech Recognition
- Developing RNN-T Models Surpassing High-Performance Hybrid Models with Customization Capability
- Internal Language Model Training for Domain-Adaptive End-to-End Speech Recognition
- Streaming End-to-End ASR based on Blockwise Non-Autoregressive Models
- Benchmarking LF-MMI, CTC and RNN-T Criteria for Streaming ASR