Sequence-Level Knowledge Distillation
arXiv:1606.07947
Abstract
Neural machine translation (NMT) offers a novel alternative formulation of translation that is potentially simpler than statistical approaches. However to reach competitive performance, NMT models need to be exceedingly large. In this paper we consider applying knowledge distillation approaches (Bucila et al., 2006; Hinton et al., 2015) that have proven successful for reducing the size of neural models in other domains to the problem of NMT. We demonstrate that standard knowledge distillation applied to word-level prediction can be effective for NMT, and also introduce two novel sequence-level versions of knowledge distillation that further improve performance, and somewhat surprisingly, seem to eliminate the need for beam search (even when applied on the original teacher model). Our best student model runs 10 times faster than its state-of-the-art teacher with little loss in performance. It is also significantly better than a baseline model trained without knowledge distillation: by 4.2/1.7 BLEU with greedy decoding/beam search. Applying weight pruning on top of knowledge distillation results in a student model that has 13 times fewer parameters than the original teacher model, with a decrease of 0.4 BLEU.
EMNLP 2016
References in corpus (9)
- Distilling the Knowledge in a Neural Network
- Sequence to Sequence Learning with Neural Networks
- Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
- Compressing Neural Networks with the Hashing Trick
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Search-based Structured Prediction
- Transferring Knowledge from a RNN to a DNN
- On the Compression of Recurrent Neural Networks with an Application to LVCSR acoustic modeling for Embedded Speech Recognition
- Learning Compact Recurrent Neural Networks
Cited by in corpus (65)
- Knowledge Distillation: A Survey
- Language Models are Few-Shot Learners
- FastSpeech: Fast, Robust and Controllable Text to Speech
- Non-Autoregressive Neural Machine Translation
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search
- Style Transfer as Unsupervised Machine Translation
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
- Soft-Label Dataset Distillation and Text Dataset Distillation
- Understanding and Improving Knowledge Distillation
- Ensemble Distillation for Neural Machine Translation
- Gram-CTC: Automatic Unit Selection and Target Decomposition for Sequence Labelling
- Theory and Experiments on Vector Quantized Autoencoders
- Domain specialization: a post-training domain adaptation for Neural Machine Translation
- Learning Approximate Inference Networks for Structured Prediction
- Domain, Translationese and Noise in Synthetic Data for Neural Machine Translation
- Incorporating BERT into Parallel Sequence Decoding with Adapters
- Imitation Learning for Non-Autoregressive Neural Machine Translation
- BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale Pretraining
- Distilling Knowledge Learned in BERT for Text Generation
- Ranking Distillation: Learning Compact Ranking Models With High Performance for Recommender System
- Detecting Hallucinated Content in Conditional Neural Sequence Generation
- Non-Autoregressive Machine Translation with Auxiliary Regularization
- Non-Autoregressive Neural Machine Translation with Enhanced Decoder Input
- Triangular Architecture for Rare Language Translation
- Token-Level Ensemble Distillation for Grapheme-to-Phoneme Conversion
- KDLSQ-BERT: A Quantized Bert Combining Knowledge Distillation with Learned Step Size Quantization
- Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning
- Semi-Autoregressive Neural Machine Translation
- Hint-Based Training for Non-Autoregressive Machine Translation
- Cascaded Text Generation with Markov Transformers
- Improved training for online end-to-end speech recognition systems
- Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation
- Dynamic Multi-Branch Layers for On-Device Neural Machine Translation
- Improving Cross-Lingual Transfer Learning for End-to-End Speech Recognition with Speech Translation
- Sharing Attention Weights for Fast Transformer
- Neural Machine Translation from Simplified Translations
- Scalable Bayesian Learning of Recurrent Neural Networks for Language Modeling
- Structural Knowledge Distillation: Tractably Distilling Information for Structured Predictor
- Adaptive Dense-to-Sparse Paradigm for Pruning Online Recommendation System with Non-Stationary Data
- MC-SF: Slow-Fast Learning for Mobile-Cloud Collaborative Recommendation
- Few-Shot NLG with Pre-Trained Language Model
- Tensor Decomposition for Compressing Recurrent Neural Network
- Cross-lingual Distillation for Text Classification
- Teaching Machines to Converse
- Power Consumption Variation over Activation Functions
- Learn Spelling from Teachers: Transferring Knowledge from Language Models to Sequence-to-Sequence Speech Recognition
- Faster Re-translation Using Non-Autoregressive Model For Simultaneous Neural Machine Translation
- Improving Non-autoregressive Neural Machine Translation with Monolingual Data
- Task-Level Curriculum Learning for Non-Autoregressive Neural Machine Translation
- VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer
- Pro-KD: Progressive Distillation by Following the Footsteps of the Teacher
- Cross-task pre-training for on-device acoustic scene classification
- Towards Variable-Length Textual Adversarial Attacks
- Recursive Top-Down Production for Sentence Generation with Latent Trees
- Towards Reinforcement Learning for Pivot-based Neural Machine Translation with Non-autoregressive Transformer
- Language-Independent Representor for Neural Machine Translation
- MTSS: Learn from Multiple Domain Teachers and Become a Multi-domain Dialogue Expert
- The USTC-NELSLIP Systems for Simultaneous Speech Translation Task at IWSLT 2021
- WeChat Neural Machine Translation Systems for WMT21
- Autoregressive Knowledge Distillation through Imitation Learning
- Subword Language Model for Query Auto-Completion
- Partial to Whole Knowledge Distillation: Progressive Distilling Decomposed Knowledge Boosts Student Better
- Hidden State Guidance: Improving Image Captioning using An Image Conditioned Autoencoder
- Learn to Talk via Proactive Knowledge Transfer