Effective Quantization Methods for Recurrent Neural Networks
arXiv:1611.10176
Abstract
Reducing bit-widths of weights, activations, and gradients of a Neural Network can shrink its storage size and memory usage, and also allow for faster training and inference by exploiting bitwise operations. However, previous attempts for quantization of RNNs show considerable performance degradation when using low bit-width weights and activations. In this paper, we propose methods to quantize the structure of gates and interlinks in LSTM and GRU cells. In addition, we propose balanced quantization methods for weights to further reduce performance degradation. Experiments on PTB and IMDB datasets confirm effectiveness of our methods as performances of our models match or surpass the previous state-of-the-art of quantized RNN.
References in corpus (6)
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Recurrent Neural Network Regularization
- Going Deeper with Convolutions
- Compressing Deep Convolutional Networks using Vector Quantization
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Recurrent Neural Networks With Limited Numerical Precision
Cited by in corpus (14)
- Knowledge Distillation from Internal Representations
- Low Precision RNNs: Quantizing RNNs Without Losing Accuracy
- Post-Training 4-bit Quantization on Embedding Tables
- Vau da muntanialas: Energy-efficient multi-die scalable acceleration of RNN inference
- Compression of Acoustic Event Detection Models with Low-rank Matrix Factorization and Quantization Training
- Precision Highway for Ultra Low-Precision Quantization
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization Framework
- Small-Footprint Open-Vocabulary Keyword Spotting with Quantized LSTM Networks
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network Quantization
- Knowledge distillation for optimization of quantized deep neural networks
- Edge Intelligence Empowered UAVs for Automated Wind Farm Monitoring in Smart Grids
- 4-bit Quantization of LSTM-based Speech Recognition Models
- Compression of Acoustic Event Detection Models With Quantized Distillation
- AutoQNN: An End-to-End Framework for Automatically Quantizing Neural Networks