Q8BERT: Quantized 8Bit BERT
arXiv:1910.06188 · doi:10.1109/EMC2-NIPS53020.2019.00016
Abstract
Recently, pre-trained Transformer based language models such as BERT and GPT, have shown great improvement in many Natural Language Processing (NLP) tasks. However, these models contain a large amount of parameters. The emergence of even larger and more accurate models such as GPT2 and Megatron, suggest a trend of large pre-trained Transformer models. However, using these large models in production environments is a complex task requiring a large amount of compute, memory and power resources. In this work we show how to perform quantization-aware training during the fine-tuning phase of BERT in order to compress BERT by with minimal accuracy loss. Furthermore, the produced quantized model can accelerate inference speed if it is optimized for 8bit Integer supporting hardware.
5 Pages, Accepted at the 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing - NeurIPS 2019
References in corpus (3)
Cited by in corpus (55)
- A Survey on Visual Transformer
- Pre-trained Models for Natural Language Processing: A Survey
- Attention Mechanism in Neural Networks: Where it Comes and Where it Goes
- VLP: A Survey on Vision-Language Pre-training
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
- The NLP Cookbook: Modern Recipes for Transformer based Deep Learning Architectures
- A Practical Survey on Faster and Lighter Transformers
- Compressing Large-Scale Transformer-Based Models: A Case Study on BERT
- TinyBERT: Distilling BERT for Natural Language Understanding
- ConvBERT: Improving BERT with Span-based Dynamic Convolution
- DynaBERT: Dynamic BERT with Adaptive Width and Depth
- On the Effect of Dropping Layers of Pre-trained Transformer Models
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
- Structured Pruning of a BERT-based Question Answering Model
- Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers
- Vision Transformer Pruning
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture Search
- BinaryBERT: Pushing the Limit of BERT Quantization
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- The Lottery Ticket Hypothesis for Pre-trained BERT Networks
- Mokey: Enabling Narrow Fixed-Point Inference for Out-of-the-Box Floating-Point Transformer Models
- Movement Pruning: Adaptive Sparsity by Fine-Tuning
- Efficient Quantized Sparse Matrix Operations on Tensor Cores
- AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search
- STI: Turbocharge NLP Inference at the Edge via Elastic Pipelining
- EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets
- A Simplified Fully Quantized Transformer for End-to-end Speech Recognition
- AutoFreeze: Automatically Freezing Model Blocks to Accelerate Fine-tuning
- I-BERT: Integer-only BERT Quantization
- BEBERT: Efficient and Robust Binary Ensemble BERT
- Composite Re-Ranking for Efficient Document Search with BERT
- KDLSQ-BERT: A Quantized Bert Combining Knowledge Distillation with Learned Step Size Quantization
- NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework
- GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference
- TernaryBERT: Distillation-aware Ultra-low Bit BERT
- DC-BERT: Decoupling Question and Document for Efficient Contextual Encoding
- EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference
- Rethinking Network Pruning -- under the Pre-train and Fine-tune Paradigm
- Mesa: A Memory-saving Training Framework for Transformers
- Compressed Deep Networks: Goodbye SVD, Hello Robust Low-Rank Approximation
- An Investigation on Different Underlying Quantization Schemes for Pre-trained Language Models
- Post-Training Quantization for Vision Transformer
- CascadeBERT: Accelerating Inference of Pre-trained Language Models via Calibrated Complete Models Cascade
- Improving Task-Agnostic BERT Distillation with Layer Mapping Search
- On the Distribution, Sparsity, and Inference-time Quantization of Attention Values in Transformers
- Adversarial Self-Supervised Data-Free Distillation for Text Classification
- Teaching a Massive Open Online Course on Natural Language Processing
- HBert + BiasCorp -- Fighting Racism on the Web
- Deep Learning Meets Projective Clustering
- Weight Squeezing: Reparameterization for Knowledge Transfer and Model Compression
- Exploring Low-Cost Transformer Model Compression for Large-Scale Commercial Reply Suggestions
- Elbert: Fast Albert with Confidence-Window Based Early Exit
- Empirical Evaluation of Deep Learning Model Compression Techniques on the WaveNet Vocoder
- Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers