A Comprehensive Study of Deep Bidirectional LSTM RNNs for Acoustic Modeling in Speech Recognition
arXiv:1606.06871 · doi:10.1109/ICASSP.2017.7952599
Abstract
We present a comprehensive study of deep bidirectional long short-term memory (LSTM) recurrent neural network (RNN) based acoustic models for automatic speech recognition (ASR). We study the effect of size and depth and train models of up to 8 layers. We investigate the training aspect and study different variants of optimization methods, batching, truncated backpropagation, different regularization techniques such as dropout and regularization, and different gradient clipping variants. The major part of the experimental analysis was performed on the Quaero corpus. Additional experiments also were performed on the Switchboard corpus. Our best LSTM model has a relative improvement in word error rate of over 14\% compared to our best feed-forward neural network (FFNN) baseline on the Quaero task. On this task, we get our best result with an 8 layer bidirectional LSTM and we show that a pretraining scheme with layer-wise construction helps for deep LSTMs. Finally we compare the training calculation time of many of the presented experiments in relation with recognition performance. All the experiments were done with RETURNN, the RWTH extensible training framework for universal recurrent neural networks in combination with RASR, the RWTH ASR toolkit.
published on ICASSP 2017 conference, New Orleans, USA
References in corpus (9)
- Deep Learning in Neural Networks: An Overview
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Improving neural networks by preventing co-adaptation of feature detectors
- ADADELTA: An Adaptive Learning Rate Method
- Theano: new features and speed improvements
- Adding Gradient Noise Improves Learning for Very Deep Networks
- Associative Long Short-Term Memory
- Deep Recurrent Neural Networks for Acoustic Modelling
- Highway Long Short-Term Memory RNNs for Distant Speech Recognition
Cited by in corpus (25)
- Improved training of end-to-end attention models for speech recognition
- An overview and comparative analysis of Recurrent Neural Networks for Short Term Load Forecasting
- Language Modeling with Deep Transformers
- RETURNN as a Generic Flexible Neural Toolkit with Application to Translation and Speech Recognition
- A Hitting Time Analysis of Stochastic Gradient Langevin Dynamics
- A New Training Pipeline for an Improved Neural Transducer
- Towards Better Domain Adaptation for Self-supervised Models: A Case Study of Child ASR
- The YouTube-8M Kaggle Competition: Challenges and Methods
- ConcealNet: An End-to-end Neural Network for Packet Loss Concealment in Deep Speech Emotion Recognition
- Acoustic Data-Driven Subword Modeling for End-to-End Speech Recognition
- On Stationary-Point Hitting Time and Ergodicity of Stochastic Gradient Langevin Dynamics
- Context-Dependent Acoustic Modeling without Explicit Phone Clustering
- Gated Recurrent Context: Softmax-free Attention for Online Encoder-Decoder Speech Recognition
- A systematic comparison of grapheme-based vs. phoneme-based label units for encoder-decoder-attention models
- Equivalence of Segmental and Neural Transducer Modeling: A Proof of Concept
- Error Reduction Network for DBLSTM-based Voice Conversion
- "I have vxxx bxx connexxxn!": Facing Packet Loss in Deep Speech Emotion Recognition
- Learning Shared Encoding Representation for End-to-End Speech Recognition Models
- Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition
- Bi-APC: Bidirectional Autoregressive Predictive Coding for Unsupervised Pre-training and Its Application to Children's ASR
- The RWTH ASR System for TED-LIUM Release 2: Improving Hybrid HMM with SpecAugment
- Using multi-task learning to improve the performance of acoustic-to-word and conventional hybrid models
- Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data
- Understanding Feature Selection and Feature Memorization in Recurrent Neural Networks
- SVGD: A Virtual Gradients Descent Method for Stochastic Optimization