Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural Networks
arXiv:1909.13144
Abstract
We propose Additive Powers-of-Two~(APoT) quantization, an efficient non-uniform quantization scheme for the bell-shaped and long-tailed distribution of weights and activations in neural networks. By constraining all quantization levels as the sum of Powers-of-Two terms, APoT quantization enjoys high computational efficiency and a good match with the distribution of weights. A simple reparameterization of the clipping function is applied to generate a better-defined gradient for learning the clipping threshold. Moreover, weight normalization is presented to refine the distribution of weights to make the training more stable and consistent. Experimental results show that our proposed method outperforms state-of-the-art methods, and is even competitive with the full-precision models, demonstrating the effectiveness of our proposed APoT quantization. For example, our 4-bit quantized ResNet-50 on ImageNet achieves 76.6% top-1 accuracy without bells and whistles; meanwhile, our model reduces 22% computational cost compared with the uniformly quantized counterpart. The code is available at https://github.com/yhhhli/APoT_Quantization.
quantization, efficient neural network
References in corpus (4)
Cited by in corpus (10)
- Training with Quantization Noise for Extreme Model Compression
- Artificial Neural Networks for Photonic Applications: From Algorithms to Implementation
- Reducing Computational Complexity of Neural Networks in Optical Channel Equalization: From Concepts to Implementation
- Elastic Significant Bit Quantization and Acceleration for Deep Neural Networks
- Robustness and Transferability of Universal Attacks on Compressed Models
- Post-training Quantization with Multiple Points: Mixed Precision without Mixed Precision
- MetaMix: Meta-state Precision Searcher for Mixed-precision Activation Quantization
- AskewSGD : An Annealed interval-constrained Optimisation method to train Quantized Neural Networks
- : Sign-Sparse-Shift Reparametrization for Effective Training of Low-bit Shift Networks
- AutoQNN: An End-to-End Framework for Automatically Quantizing Neural Networks