Oracle Teacher: Leveraging Target Information for Better Knowledge Distillation of CTC Models
arXiv:2111.03664 · doi:10.1109/TASLP.2023.3297955
Abstract
Knowledge distillation (KD), best known as an effective method for model compression, aims at transferring the knowledge of a bigger network (teacher) to a much smaller network (student). Conventional KD methods usually employ the teacher model trained in a supervised manner, where output labels are treated only as targets. Extending this supervised scheme further, we introduce a new type of teacher model for connectionist temporal classification (CTC)-based sequence models, namely Oracle Teacher, that leverages both the source inputs and the output labels as the teacher model's input. Since the Oracle Teacher learns a more accurate CTC alignment by referring to the target information, it can provide the student with more optimal guidance. One potential risk for the proposed approach is a trivial solution that the model's output directly copies the target input. Based on a many-to-one mapping property of the CTC algorithm, we present a training strategy that can effectively prevent the trivial solution and thus enables utilizing both source and target inputs for model training. Extensive experiments are conducted on two sequence learning tasks: speech recognition and scene text recognition. From the experimental results, we empirically show that the proposed model improves the students across these tasks while achieving a considerable speed-up in the teacher model's training time.
Accepted by IEEE/ACM Transactions on Audio, Speech and Language Processing
References in corpus (10)
- Distilling the Knowledge in a Neural Network
- ADADELTA: An Adaptive Learning Rate Method
- FitNets: Hints for Thin Deep Nets
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- Sequence Transduction with Recurrent Neural Networks
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- Rosetta: Large scale system for text detection and recognition in images
- Efficiently Fusing Pretrained Acoustic and Linguistic Encoders for Low-resource Speech Recognition
- Non-autoregressive Transformer-based End-to-end ASR using BERT
- TutorNet: Towards Flexible Knowledge Distillation for End-to-End Speech Recognition