Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search
arXiv:2010.07003
Abstract
Despite transformers' impressive accuracy, their computational cost is often prohibitive to use with limited computational resources. Most previous approaches to improve inference efficiency require a separate model for each possible computational budget. In this paper, we extend PoWER-BERT (Goyal et al., 2020) and propose Length-Adaptive Transformer that can be used for various inference scenarios after one-shot training. We train a transformer with LengthDrop, a structural variant of dropout, which stochastically determines a sequence length at each layer. We then conduct a multi-objective evolutionary search to find a length configuration that maximizes the accuracy and minimizes the efficiency metric under any given computational budget. Additionally, we significantly extend the applicability of PoWER-BERT beyond sequence-level classification into token-level classification with Drop-and-Restore process that drops word-vectors temporarily in intermediate layers and restores at the last layer if necessary. We empirically verify the utility of the proposed approach by demonstrating the superior accuracy-efficiency trade-off under various setups, including span-based question answering and text classification. Code is available at https://github.com/clovaai/length-adaptive-transformer.
ACL 2021; 11 pages, 4 figures
References in corpus (16)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Distilling the Knowledge in a Neural Network
- Language Models are Few-Shot Learners
- Scaling Laws for Neural Language Models
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Reducing Transformer Depth on Demand with Structured Dropout
- Slimmable Neural Networks
- Long Range Arena: A Benchmark for Efficient Transformers
- Rethinking Attention with Performers
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing
- NSML: A Machine Learning Platform That Enables You to Focus on Your Models
- Depth-Adaptive Transformer
- FastBERT: a Self-distilling BERT with Adaptive Inference Time
- TR-BERT: Dynamic Token Reduction for Accelerating BERT Inference
- Accelerating BERT Inference for Sequence Labeling via Early-Exit
- Towards Accurate and Reliable Energy Measurement of NLP Models