Compressing Large-Scale Transformer-Based Models: A Case Study on BERT
arXiv:2002.11985 · doi:10.1162/tacl_a_00413
Abstract
Pre-trained Transformer-based models have achieved state-of-the-art performance for various Natural Language Processing (NLP) tasks. However, these models often have billions of parameters, and, thus, are too resource-hungry and computation-intensive to suit low-capability devices or applications with strict latency requirements. One potential remedy for this is model compression, which has attracted a lot of research attention. Here, we summarize the research in compressing Transformers, focusing on the especially popular BERT model. In particular, we survey the state of the art in compression for BERT, we clarify the current best practices for compressing large-scale Transformer models, and we provide insights into the workings of various methods. Our categorization and analysis also shed light on promising future research directions for achieving lightweight, accurate, and generic NLP models.
To appear in TACL 2021. The arXiv version is a pre-MIT Press publication version
References in corpus (4)
Cited by in corpus (12)
- Transformers in Healthcare: A Survey
- Greedy-layer Pruning: Speeding up Transformer Models for Natural Language Processing
- Mokey: Enabling Narrow Fixed-Point Inference for Out-of-the-Box Floating-Point Transformer Models
- Efficient Fine-Tuning of BERT Models on the Edge
- Investigating Hallucinations in Pruned Large Language Models for Abstractive Summarization
- CANAL -- Cyber Activity News Alerting Language Model: Empirical Approach vs. Expensive LLM
- ALPINE: An adaptive language-agnostic pruning method for language models for code
- Can persistent homology whiten Transformer-based black-box models? A case study on BERT compression
- Classification of integers based on residue classes via modern deep learning algorithms
- FANAL -- Financial Activity News Alerting Language Modeling Framework
- Transformer Explainer: Learning LLM Transformers with Interactive Visual Explanation and Experimentation
- A Runtime-Adaptive Transformer Neural Network Accelerator on FPGAs