When BERT Plays the Lottery, All Tickets Are Winning
arXiv:2005.00561
Abstract
Large Transformer-based models were shown to be reducible to a smaller number of self-attention heads and layers. We consider this phenomenon from the perspective of the lottery ticket hypothesis, using both structured and magnitude pruning. For fine-tuned BERT, we show that (a) it is possible to find subnetworks achieving performance that is comparable with that of the full model, and (b) similarly-sized subnetworks sampled from the rest of the model perform worse. Strikingly, with structured pruning even the worst possible subnetworks remain highly trainable, indicating that most pre-trained BERT weights are potentially useful. We also study the "good" subnetworks to see if their success can be attributed to superior linguistic knowledge, but find them unstable, and not explained by meaningful self-attention patterns.
EMNLP 2020 camera-ready
References in corpus (3)
Cited by in corpus (14)
- A Survey on Visual Transformer
- The Lottery Ticket Hypothesis for Pre-trained BERT Networks
- The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models
- EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets
- Pre-Trained Models: Past, Present and Future
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition
- Sanity Checks for Lottery Tickets: Does Your Winning Ticket Really Win the Jackpot?
- Drawing Robust Scratch Tickets: Subnetworks with Inborn Robustness Are Found within Randomly Initialized Networks
- SuperShaper: Task-Agnostic Super Pre-training of BERT Models with Variable Hidden Dimensions
- Super Tickets in Pre-Trained Language Models: From Model Compression to Improving Generalization
- Spending Your Winning Lottery Better After Drawing It
- Why Can You Lay Off Heads? Investigating How BERT Heads Transfer
- Understanding and Overcoming the Challenges of Efficient Transformer Quantization
- Influence Patterns for Explaining Information Flow in BERT