Reducing Transformer Depth on Demand with Structured Dropout
arXiv:1909.11556
Abstract
Overparameterized transformer networks have obtained state of the art results in various natural language processing tasks, such as machine translation, language modeling, and question answering. These models contain hundreds of millions of parameters, necessitating a large amount of computation and making them prone to overfitting. In this work, we explore LayerDrop, a form of structured dropout, which has a regularization effect during training and allows for efficient pruning at inference time. In particular, we show that it is possible to select sub-networks of any depth from one large network without having to finetune them and with limited impact on performance. We demonstrate the effectiveness of our approach by improving the state of the art on machine translation, language modeling, summarization, question answering, and language understanding benchmarks. Moreover, we show that our approach leads to small BERT-like models of higher quality compared to training from scratch or using distillation.
References in corpus (19)
- Improving neural networks by preventing co-adaptation of feature detectors
- Teaching Machines to Read and Comprehend
- XLNet: Generalized Autoregressive Pretraining for Language Understanding
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Cross-lingual Language Model Pretraining
- Pruning Filters for Efficient ConvNets
- Pointer Sentinel Mixture Models
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
- Pay Less Attention with Lightweight and Dynamic Convolutions
- Fixup Initialization: Residual Learning Without Normalization
- Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
- Adaptive Input Representations for Neural Language Modeling
- Are Sixteen Heads Really Better than One?
- Controllable Abstractive Summarization
- Pre-trained Language Model Representations for Language Generation
- Adaptively Sparse Transformers
- Auto-Sizing Neural Networks: With Applications to n-gram Language Models
- ELI5: Long Form Question Answering
Cited by in corpus (73)
- A Survey on Visual Transformer
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Beyond English-Centric Multilingual Machine Translation
- Filter-enhanced MLP is All You Need for Sequential Recommendation
- MLS: A Large-Scale Multilingual Dataset for Speech Research
- VLP: A Survey on Vision-Language Pre-training
- The NLP Cookbook: Modern Recipes for Transformer based Deep Learning Architectures
- End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
- TinyBERT: Distilling BERT for Natural Language Understanding
- ConvBERT: Improving BERT with Span-based Dynamic Convolution
- DynaBERT: Dynamic BERT with Adaptive Width and Depth
- PyTorch Distributed: Experiences on Accelerating Data Parallel Training
- Training with Quantization Noise for Extreme Model Compression
- Fine-tuning wav2vec2 for speaker recognition
- TransReID: Transformer-based Object Re-Identification
- On the Effect of Dropping Layers of Pre-trained Transformer Models
- Representation Learning for Natural Language Processing
- Long-Range Transformers for Dynamic Spatiotemporal Forecasting
- Structured Pruning of a BERT-based Question Answering Model
- FlauBERT: Unsupervised Language Model Pre-training for French
- Pre-trained Summarization Distillation
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture Search
- BinaryBERT: Pushing the Limit of BERT Quantization
- BERT Loses Patience: Fast and Robust Inference with Early Exit
- Only Train Once: A One-Shot Neural Network Training And Pruning Framework
- The Lottery Ticket Hypothesis for Pre-trained BERT Networks
- Greedy-layer Pruning: Speeding up Transformer Models for Natural Language Processing
- BERT-of-Theseus: Compressing BERT by Progressive Module Replacing
- Movement Pruning: Adaptive Sparsity by Fine-Tuning
- Addressing Some Limitations of Transformers with Feedback Memory
- DDPNAS: Efficient Neural Architecture Search via Dynamic Distribution Pruning
- Deep Transformers with Latent Depth
- On Optimal Transformer Depth for Low-Resource Language Translation
- Differentiable Model Compression via Pseudo Quantization Noise
- DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference
- Contrastive Visual-Linguistic Pretraining
- EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets
- PowerNorm: Rethinking Batch Normalization in Transformers
- AutoFreeze: Automatically Freezing Model Blocks to Accelerate Fine-tuning
- I-BERT: Integer-only BERT Quantization
- Pre-Trained Models: Past, Present and Future
- The EarlyBIRD Catches the Bug: On Exploiting Early Layers of Encoder Models for More Efficient Code Classification
- KDLSQ-BERT: A Quantized Bert Combining Knowledge Distillation with Learned Step Size Quantization
- Controlling Computation versus Quality for Neural Sequence Models
- GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference
- Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
- TernaryBERT: Distillation-aware Ultra-low Bit BERT
- An Efficient Transformer Decoder with Compressed Sub-layers
- Training speaker recognition systems with limited data
- Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search
- BERT's output layer recognizes all hidden layers? Some Intriguing Phenomena and a simple way to boost BERT
- Alleviating the Inequality of Attention Heads for Neural Machine Translation
- Scheduled DropHead: A Regularization Method for Transformer Models
- Rethinking Network Pruning -- under the Pre-train and Fine-tune Paradigm
- A Unified Pruning Framework for Vision Transformers
- Compressed Deep Networks: Goodbye SVD, Hello Robust Low-Rank Approximation
- CascadeBERT: Accelerating Inference of Pre-trained Language Models via Calibrated Complete Models Cascade
- SuperShaper: Task-Agnostic Super Pre-training of BERT Models with Variable Hidden Dimensions
- Training Flexible Depth Model by Multi-Task Learning for Neural Machine Translation
- Adversarial Self-Supervised Data-Free Distillation for Text Classification
- Exceeding the Limits of Visual-Linguistic Multi-Task Learning
- Super Tickets in Pre-Trained Language Models: From Model Compression to Improving Generalization
- Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency
- Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers
- Generating Diverse Translation from Model Distribution with Dropout
- Deep Learning Meets Projective Clustering
- Population Based Training for Data Augmentation and Regularization in Speech Recognition
- Weight Squeezing: Reparameterization for Knowledge Transfer and Model Compression
- Consistent Accelerated Inference via Confident Adaptive Transformers
- LightSeq2: Accelerated Training for Transformer-based Models on GPUs
- A Survey on Green Deep Learning
- Pruning Attention Heads of Transformer Models Using A* Search: A Novel Approach to Compress Big NLP Architectures
- When in Doubt, Summon the Titans: Efficient Inference with Large Models