BERT-of-Theseus: Compressing BERT by Progressive Module Replacing
arXiv:2002.02925
Abstract
In this paper, we propose a novel model compression approach to effectively compress BERT by progressive module replacing. Our approach first divides the original BERT into several modules and builds their compact substitutes. Then, we randomly replace the original modules with their substitutes to train the compact modules to mimic the behavior of the original modules. We progressively increase the probability of replacement through the training. In this way, our approach brings a deeper level of interaction between the original and compact models. Compared to the previous knowledge distillation approaches for BERT compression, our approach does not introduce any additional loss function. Our approach outperforms existing knowledge distillation approaches on GLUE benchmark, showing a new perspective of model compression.
EMNLP 2020
References in corpus (9)
- Distilling the Knowledge in a Neural Network
- Compressing Deep Convolutional Networks using Vector Quantization
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
- Reformer: The Efficient Transformer
- Reducing Transformer Depth on Demand with Structured Dropout
- Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
- Multilingual Neural Machine Translation with Knowledge Distillation
Cited by in corpus (10)
- Pre-trained Models for Natural Language Processing: A Survey
- Compressing Large-Scale Transformer-Based Models: A Case Study on BERT
- Pre-trained Summarization Distillation
- Optimal Subarchitecture Extraction For BERT
- MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers
- NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework
- Not All Attention Is All You Need
- KroneckerBERT: Learning Kronecker Decomposition for Pre-trained Language Models via Knowledge Distillation
- Which *BERT? A Survey Organizing Contextualized Encoders
- SuperShaper: Task-Agnostic Super Pre-training of BERT Models with Variable Hidden Dimensions