Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
arXiv:1908.08962
Abstract
Recent developments in natural language representations have been accompanied by large and expensive models that leverage vast amounts of general-domain text through self-supervised pre-training. Due to the cost of applying such models to down-stream tasks, several model compression techniques on pre-trained language representations have been proposed (Sun et al., 2019; Sanh, 2019). However, surprisingly, the simple baseline of just pre-training and fine-tuning compact models has been overlooked. In this paper, we first show that pre-training remains important in the context of smaller architectures, and fine-tuning pre-trained compact models can be competitive to more elaborate methods proposed in concurrent work. Starting with pre-trained compact models, we then explore transferring task knowledge from large fine-tuned models through standard knowledge distillation. The resulting simple, yet effective and general algorithm, Pre-trained Distillation, brings further improvements. Through extensive experiments, we more generally explore the interaction between pre-training and distillation under two variables that have been under-studied: model size and properties of unlabeled task data. One surprising observation is that they have a compound effect even when sequentially applied on the same data. To accelerate future research, we will make our 24 pre-trained miniature BERT models publicly available.
Added comparison to concurrent work
References in corpus (4)
Cited by in corpus (58)
- Reducing Transformer Depth on Demand with Structured Dropout
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
- SEED: Self-supervised Distillation For Visual Representation
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture Search
- Learning from others' mistakes: Avoiding dataset biases without modeling them
- Zero-shot Node Classification with Decomposed Graph Prototype Network
- MixKD: Towards Efficient Distillation of Large-scale Language Models
- End-to-End Self-Debiasing Framework for Robust NLU Training
- TwinBERT: Distilling Knowledge to Twin-Structured BERT Models for Efficient Retrieval
- Optimal Subarchitecture Extraction For BERT
- Distilling Large Language Models into Tiny and Effective Students using pQRNN
- Accenture at CheckThat! 2020: If you say so: Post-hoc fact-checking of claims using transformer-based models
- Tiny Transformers for Environmental Sound Classification at the Edge
- You Only Compress Once: Towards Effective and Elastic BERT Compression via Exploit-Explore Stochastic Nature Gradient
- XtremeDistilTransformers: Task Transfer for Task-agnostic Distillation
- Sentence Embeddings using Supervised Contrastive Learning
- Large-Scale News Classification using BERT Language Model: Spark NLP Approach
- Multi-stage Progressive Compression of Conformer Transducer for On-device Speech Recognition
- Know What You Don't Need: Single-Shot Meta-Pruning for Attention Heads
- LightMBERT: A Simple Yet Effective Method for Multilingual BERT Distillation
- LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding
- FrugalScore: Learning Cheaper, Lighter and Faster Evaluation Metricsfor Automatic Text Generation
- Efficient pre-training objectives for Transformers
- Grounding Representation Similarity with Statistical Testing
- Demystifying BERT: Implications for Accelerator Design
- IrEne: Interpretable Energy Prediction for Transformers
- Fine-grained Sentiment Controlled Text Generation
- Is neural language acquisition similar to natural? A chronological probing study
- ViBERTgrid: A Jointly Trained Multi-Modal 2D Document Representation for Key Information Extraction from Documents
- Deep Neural Compression Via Concurrent Pruning and Self-Distillation
- SuperShaper: Task-Agnostic Super Pre-training of BERT Models with Variable Hidden Dimensions
- Follow Your Path: a Progressive Method for Knowledge Distillation
- Improving Task-Agnostic BERT Distillation with Layer Mapping Search
- Weakly-Supervised Open-Retrieval Conversational Question Answering
- Towards Structured Dynamic Sparse Pre-Training of BERT
- nmT5 -- Is parallel data still relevant for pre-training massively multilingual language models?
- Retraining DistilBERT for a Voice Shopping Assistant by Using Universal Dependencies
- Lookup or Exploratory: What is Your Search Intent?
- Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color
- PAUSE: Positive and Annealed Unlabeled Sentence Embedding
- Accenture at CheckThat! 2021: Interesting claim identification and ranking with contextually sensitive lexical training data augmentation
- MergeDistill: Merging Pre-trained Language Models using Distillation
- Characterizing Abhorrent, Misinformative, and Mistargeted Content on YouTube
- Programming with Neural Surrogates of Programs
- Debiasing Methods in Natural Language Understanding Make Bias More Accessible
- Why Can You Lay Off Heads? Investigating How BERT Heads Transfer
- Incorporating Residual and Normalization Layers into Analysis of Masked Language Models
- EfficientBERT: Progressively Searching Multilayer Perceptron via Warm-up Knowledge Distillation
- Structural analysis of an all-purpose question answering model
- Distilling Linguistic Context for Language Model Compression
- Knowledge Distillation with Noisy Labels for Natural Language Understanding
- Distiller: A Systematic Study of Model Distillation Methods in Natural Language Processing
- Generalization in NLI: Ways (Not) To Go Beyond Simple Heuristics
- Context-Aware Transformer Transducer for Speech Recognition
- AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language Models
- Frustratingly Simple Pretraining Alternatives to Masked Language Modeling
- DNN-Based Semantic Model for Rescoring N-best Speech Recognition List
- Coreference Augmentation for Multi-Domain Task-Oriented Dialogue State Tracking