Hierarchical Multitask Learning for CTC-based Speech Recognition
arXiv:1807.06234
Abstract
Previous work has shown that neural encoder-decoder speech recognition can be improved with hierarchical multitask learning, where auxiliary tasks are added at intermediate layers of a deep encoder. We explore the effect of hierarchical multitask learning in the context of connectionist temporal classification (CTC)-based speech recognition, and investigate several aspects of this approach. Consistent with previous work, we observe performance improvements on telephone conversational speech recognition (specifically the Eval2000 test sets) when training a subword-level CTC model with an auxiliary phone loss at an intermediate layer. We analyze the effects of a number of experimental variables (like interpolation constant and position of the auxiliary loss function), performance in lower-resource settings, and the relationship between pretraining and multitask learning. We observe that the hierarchical multitask approach improves over standard multitask training in our higher-data experiments, while in the low-resource settings standard multitask training works well. The best results are obtained by combining hierarchical multitask learning and pretraining, which improves word error rates by 3.4% absolute on the Eval2000 test sets.
Technical Report
References in corpus (5)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Recurrent Neural Network Regularization
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
- Improved training of end-to-end attention models for speech recognition
- Advances in All-Neural Speech Recognition
Cited by in corpus (7)
- OmniNet: A unified architecture for multi-modal multi-task learning
- Multiresolution and Multimodal Speech Recognition with Transformers
- Intermediate Loss Regularization for CTC-based Speech Recognition
- Leveraging Hierarchical Structures for Few-Shot Musical Instrument Recognition
- On the Inductive Bias of Word-Character-Level Multi-Task Learning for Speech Recognition
- Learning acoustic word embeddings with phonetically associated triplet network
- Acoustically Grounded Word Embeddings for Improved Acoustics-to-Word Speech Recognition