Recent Progresses in Deep Learning based Acoustic Models (Updated)
arXiv:1804.09298
Abstract
In this paper, we summarize recent progresses made in deep learning based acoustic models and the motivation and insights behind the surveyed techniques. We first discuss acoustic models that can effectively exploit variable-length contextual information, such as recurrent neural networks (RNNs), convolutional neural networks (CNNs), and their various combination with other models. We then describe acoustic models that are optimized end-to-end with emphasis on feature representations learned jointly with rest of the system, the connectionist temporal classification (CTC) criterion, and the attention-based sequence-to-sequence model. We further illustrate robustness issues in speech recognition systems, and discuss acoustic model adaptation, speech enhancement and separation, and robust training strategies. We also cover modeling techniques that lead to more efficient decoding and discuss possible future directions in acoustic model research.
This is an updated version with latest literature until ICASSP2018 of the paper: Dong Yu and Jinyu Li, "Recent Progresses in Deep Learning based Acoustic Models," vol.4, no.3, IEEE/CAA Journal of Automatica Sinica, 2017
References in corpus (15)
- Distilling the Knowledge in a Neural Network
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Speaker-Invariant Training via Adversarial Learning
- Adversarial Teacher-Student Learning for Unsupervised Domain Adaptation
- Exploring Neural Transducers for End-to-End Speech Recognition
- Unsupervised Adaptation with Domain Separation Networks for Robust Speech Recognition
- Invariant Representations for Noisy Speech Recognition
- Neural Speech Recognizer: Acoustic-to-Word LSTM Model for Large Vocabulary Speech Recognition
- Transferring Knowledge from a RNN to a DNN
- Dense Prediction on Sequences with Time-Dilated Convolutions for Speech Recognition
- Direct Acoustics-to-Word Models for English Conversational Speech Recognition
- Large-Scale Domain Adaptation via Teacher-Student Learning
- Domain Adversarial Training for Accented Speech Recognition
- Boosting Noise Robustness of Acoustic Model via Deep Adversarial Training
- On the efficient representation and execution of deep acoustic models