Learning Hidden Unit Contributions for Unsupervised Acoustic Model Adaptation
arXiv:1601.02828 · doi:10.1109/TASLP.2016.2560534
Abstract
This work presents a broad study on the adaptation of neural network acoustic models by means of learning hidden unit contributions (LHUC) -- a method that linearly re-combines hidden units in a speaker- or environment-dependent manner using small amounts of unsupervised adaptation data. We also extend LHUC to a speaker adaptive training (SAT) framework that leads to a more adaptable DNN acoustic model, working both in a speaker-dependent and a speaker-independent manner, without the requirements to maintain auxiliary speaker-dependent feature extractors or to introduce significant speaker-dependent changes to the DNN structure. Through a series of experiments on four different speech recognition benchmarks (TED talks, Switchboard, AMI meetings, and Aurora4) comprising 270 test speakers, we show that LHUC in both its test-only and SAT variants results in consistent word error rate reductions ranging from 5% to 23% relative depending on the task and the degree of mismatch between training and test data. In addition, we have investigated the effect of the amount of adaptation data per speaker, the quality of unsupervised adaptation targets, the complementarity to other adaptation techniques, one-shot adaptation, and an extension to adapting DNNs trained in a sequence discriminative manner.
14 pages, 9 Tables, 11 Figues in IEEE/ACM Transactions on Audio, Speech and Language Processing, Vol. 24, Num. 8, 2016
Cited by in corpus (30)
- Speaker-Invariant Training via Adversarial Learning
- Conditional Teacher-Student Learning
- Adversarial Teacher-Student Learning for Unsupervised Domain Adaptation
- Adaptation Algorithms for Neural Network-Based Speech Recognition: An Overview
- Unsupervised Adaptation with Domain Separation Networks for Robust Speech Recognition
- Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends
- Investigation of Data Augmentation Techniques for Disordered Speech Recognition
- Speaker Adaptation for Attention-Based End-to-End Speech Recognition
- Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition
- From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition
- Bayesian Learning for Deep Neural Network Adaptation
- Spectro-Temporal Deep Features for Disordered Speech Assessment and Recognition
- Small-footprint Highway Deep Neural Networks for Speech Recognition
- Very Deep Convolutional Neural Networks for Robust Speech Recognition
- Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
- Dynamic Layer Normalization for Adaptive Neural Acoustic Modeling in Speech Recognition
- Differentiable Pooling for Unsupervised Acoustic Model Adaptation
- Generalizing Speaker Verification for Spoof Awareness in the Embedding Space
- Multilingual and Unsupervised Subword Modeling for Zero-Resource Languages
- Lattice-Based Unsupervised Test-Time Adaptation of Neural Network Acoustic Models
- Two-pass Decoding and Cross-adaptation Based System Combination of End-to-end Conformer and Hybrid TDNN ASR Systems
- Recent Progresses in Deep Learning based Acoustic Models (Updated)
- Adversarial Data Augmentation for Disordered Speech Recognition
- TS-Net: OCR Trained to Switch Between Text Transcription Styles
- Exploring Gaussian mixture model framework for speaker adaptation of deep neural network acoustic models
- Residual Adapters for Parameter-Efficient ASR Adaptation to Atypical and Accented Speech
- A Unified Speaker Adaptation Approach for ASR
- Attention-based gated scaling adaptative acoustic model for ctc-based speech recognition
- Speaker Adaptation for End-to-End CTC Models
- Semi-tied Units for Efficient Gating in LSTM and Highway Networks