Speaker Adaptation for Attention-Based End-to-End Speech Recognition
arXiv:1911.03762 · doi:10.21437/Interspeech.2019-3135
Abstract
We propose three regularization-based speaker adaptation approaches to adapt the attention-based encoder-decoder (AED) model with very limited adaptation data from target speakers for end-to-end automatic speech recognition. The first method is Kullback-Leibler divergence (KLD) regularization, in which the output distribution of a speaker-dependent (SD) AED is forced to be close to that of the speaker-independent (SI) model by adding a KLD regularization to the adaptation criterion. To compensate for the asymmetric deficiency in KLD regularization, an adversarial speaker adaptation (ASA) method is proposed to regularize the deep-feature distribution of the SD AED through the adversarial learning of an auxiliary discriminator and the SD AED. The third approach is the multi-task learning, in which an SD AED is trained to jointly perform the primary task of predicting a large number of output units and an auxiliary task of predicting a small number of output units to alleviate the target sparsity issue. Evaluated on a Microsoft short message dictation task, all three methods are highly effective in adapting the AED model, achieving up to 12.2% and 3.0% word error rate improvement over an SI AED trained from 3400 hours data for supervised and unsupervised adaptation, respectively.
5 pages, 3 figures, Interspeech 2019
References in corpus (18)
- Neural Machine Translation by Jointly Learning to Align and Translate
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Unsupervised Domain Adaptation by Backpropagation
- Attention-Based Models for Speech Recognition
- Sequence Transduction with Recurrent Neural Networks
- Learning Hidden Unit Contributions for Unsupervised Acoustic Model Adaptation
- Speaker-Invariant Training via Adversarial Learning
- Conditional Teacher-Student Learning
- Deep Long Short-Term Memory Adaptive Beamforming Networks For Multichannel Robust Speech Recognition
- Adversarial Teacher-Student Learning for Unsupervised Domain Adaptation
- Unsupervised Adaptation with Domain Separation Networks for Robust Speech Recognition
- Invariant Representations for Noisy Speech Recognition
- Adversarial Speaker Verification
- Cycle-Consistent Speech Enhancement
- Adversarial Feature-Mapping for Speech Enhancement
- Adversarial Speaker Adaptation
- Maximum a Posteriori Adaptation of Network Parameters in Deep Models
- Attentive Adversarial Learning for Domain-Invariant Training
Cited by in corpus (9)
- Adaptation Algorithms for Neural Network-Based Speech Recognition: An Overview
- Bayesian Learning for Deep Neural Network Adaptation
- Unsupervised Speaker Adaptation using Attention-based Speaker Memory for End-to-End ASR
- L-Vector: Neural Label Embedding for Domain Adaptation
- Internal Language Model Training for Domain-Adaptive End-to-End Speech Recognition
- Speaker Re-identification with Speaker Dependent Speech Enhancement
- Generative Adversarial Training Data Adaptation for Very Low-resource Automatic Speech Recognition
- Multiple-hypothesis CTC-based semi-supervised adaptation of end-to-end speech recognition
- Character-Aware Attention-Based End-to-End Speech Recognition