Adaptation Algorithms for Neural Network-Based Speech Recognition: An Overview
arXiv:2008.06580 · doi:10.1109/OJSP.2020.3045349
Abstract
We present a structured overview of adaptation algorithms for neural network-based speech recognition, considering both hybrid hidden Markov model / neural network systems and end-to-end neural network systems, with a focus on speaker adaptation, domain adaptation, and accent adaptation. The overview characterizes adaptation algorithms as based on embeddings, model parameter adaptation, or data augmentation. We present a meta-analysis of the performance of speech recognition adaptation algorithms, based on relative error rate reductions as reported in the literature.
Total of 31 pages, 27 figures. Associated repository: https://github.com/pswietojanski/ojsp_adaptation_review_2020
References in corpus (25)
- Distilling the Knowledge in a Neural Network
- Unsupervised Domain Adaptation by Backpropagation
- Attention-Based Models for Speech Recognition
- Sequence Transduction with Recurrent Neural Networks
- On Using Monolingual Corpora in Neural Machine Translation
- First-Pass Large Vocabulary Continuous Speech Recognition using Bi-Directional Recurrent DNNs
- Conditional Teacher-Student Learning
- Improving Neural Language Models with a Continuous Cache
- Exploring Neural Transducers for End-to-End Speech Recognition
- Transformer-Transducer: End-to-End Speech Recognition with Self-Attention
- Dynamic Evaluation of Neural Sequence Models
- Speaker Adaptation for Attention-Based End-to-End Speech Recognition
- Adversarial Speaker Adaptation
- A Highly Adaptive Acoustic Model for Accurate Multi-Dialect Speech Recognition
- Leveraging native language information for improved accented speech recognition
- Cross-Language Transfer Learning, Continuous Learning, and Domain Adaptation for End-to-End Automatic Speech Recognition
- Semi-Supervised Model Training for Unbounded Conversational Speech Recognition
- Dynamic Layer Normalization for Adaptive Neural Acoustic Modeling in Speech Recognition
- Multi-Dialect Speech Recognition With A Single Sequence-To-Sequence Model
- Lattice-Based Unsupervised Test-Time Adaptation of Neural Network Acoustic Models
- Embedding-Based Speaker Adaptive Training of Deep Neural Networks
- DARTS: Dialectal Arabic Transcription System
- Improving RNN Transducer Modeling for End-to-End Speech Recognition
- Listen, Attend, Spell and Adapt: Speaker Adapted Sequence-to-Sequence ASR
- Personalization of End-to-end Speech Recognition On Mobile Devices For Named Entities
Cited by in corpus (8)
- Self-Supervised Speech Representation Learning: A Review
- Generalizing Speaker Verification for Spoof Awareness in the Embedding Space
- A Simple Baseline for Domain Adaptation in End to End ASR Systems Using Synthetic Data
- TS-Net: OCR Trained to Switch Between Text Transcription Styles
- On Addressing Practical Challenges for RNN-Transducer
- Minimum Word Error Rate Training with Language Model Fusion for End-to-End Speech Recognition
- Seed Words Based Data Selection for Language Model Adaptation
- Efficient Extraction of Noise-Robust Discrete Units from Self-Supervised Speech Models