Structured Pruning of a BERT-based Question Answering Model
arXiv:1910.06360
Abstract
The recent trend in industry-setting Natural Language Processing (NLP) research has been to operate large %scale pretrained language models like BERT under strict computational limits. While most model compression work has focused on "distilling" a general-purpose language representation using expensive pretraining distillation, less attention has been paid to creating smaller task-specific language representations which, arguably, are more useful in an industry setting. In this paper, we investigate compressing BERT- and RoBERTa-based question answering systems by structured pruning of parameters from the underlying transformer model. We find that an inexpensive combination of task-specific structured pruning and task-specific distillation, without the expense of pretraining distillation, yields highly-performing models across a range of speed/accuracy tradeoff operating points. We start from existing full-size models trained for SQuAD 2.0 or Natural Questions and introduce gates that allow selected parts of transformers to be individually eliminated. Specifically, we investigate (1) structured pruning to reduce the number of parameters in each transformer layer, (2) applicability to both BERT- and RoBERTa-based models, (3) applicability to both SQuAD 2.0 and Natural Questions, and (4) combining structured pruning with distillation. We achieve a near-doubling of inference speed with less than a 0.5 F1-point loss in short answer accuracy on Natural Questions.
References in corpus (14)
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
- Q8BERT: Quantized 8Bit BERT
- The State of Sparsity in Deep Neural Networks
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
- Reducing Transformer Depth on Demand with Structured Dropout
- TinyBERT: Distilling BERT for Natural Language Understanding
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
- Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
- Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers
- Model Compression with Multi-Task Knowledge Distillation for Web-scale Question Answering System
- Small and Practical BERT Models for Sequence Labeling
- DeFormer: Decomposing Pre-trained Transformers for Faster Question Answering
Cited by in corpus (26)
- On the Opportunities and Risks of Foundation Models
- TinyBERT: Distilling BERT for Natural Language Understanding
- DynaBERT: Dynamic BERT with Adaptive Width and Depth
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture Search
- When BERT Plays the Lottery, All Tickets Are Winning
- The Lottery Ticket Hypothesis for Pre-trained BERT Networks
- A Programmable Approach to Neural Network Compression
- Towards Efficient Post-training Quantization of Pre-trained Language Models
- EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets
- AutoFreeze: Automatically Freezing Model Blocks to Accelerate Fine-tuning
- KDLSQ-BERT: A Quantized Bert Combining Knowledge Distillation with Learned Step Size Quantization
- AUBER: Automated BERT Regularization
- Rethinking Network Pruning -- under the Pre-train and Fine-tune Paradigm
- Know What You Don't Need: Single-Shot Meta-Pruning for Attention Heads
- Compressed Deep Networks: Goodbye SVD, Hello Robust Low-Rank Approximation
- Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation
- Distilling Knowledge from Pre-trained Language Models via Text Smoothing
- Improving Task-Agnostic BERT Distillation with Layer Mapping Search
- Knowledge Transfer via Pre-training for Recommendation: A Review and Prospect
- DSEE: Dually Sparsity-embedded Efficient Tuning of Pre-trained Language Models
- Adversarial Self-Supervised Data-Free Distillation for Text Classification
- A Survey on Green Deep Learning
- Structural analysis of an all-purpose question answering model
- Weight Squeezing: Reparameterization for Knowledge Transfer and Model Compression
- Deep Learning Meets Projective Clustering