Towards the Limit of Network Quantization
arXiv:1612.01543
Abstract
Network quantization is one of network compression techniques to reduce the redundancy of deep neural networks. It reduces the number of distinct network parameter values by quantization in order to save the storage for them. In this paper, we design network quantization schemes that minimize the performance loss due to quantization given a compression ratio constraint. We analyze the quantitative relation of quantization errors to the neural network loss function and identify that the Hessian-weighted distortion measure is locally the right objective function for the optimization of network quantization. As a result, Hessian-weighted k-means clustering is proposed for clustering network parameters to quantize. When optimal variable-length binary codes, e.g., Huffman codes, are employed for further compression, we derive that the network quantization problem can be related to the entropy-constrained scalar quantization (ECSQ) problem in information theory and consequently propose two solutions of ECSQ for network quantization, i.e., uniform quantization and an iterative solution similar to Lloyd's algorithm. Finally, using the simple uniform quantization followed by Huffman coding, we show from our experiments that the compression ratios of 51.25, 22.17 and 40.65 are achievable for LeNet, 32-layer ResNet and AlexNet, respectively.
Published as a conference paper at ICLR 2017
References in corpus (4)
Cited by in corpus (28)
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations
- A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects
- A Survey on Methods and Theories of Quantized Neural Networks
- DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks
- And the Bit Goes Down: Revisiting the Quantization of Neural Networks
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Quantization for Rapid Deployment of Deep Neural Networks
- Fixed-point Quantization of Convolutional Neural Networks for Quantized Inference on Embedded Platforms
- Rate Distortion For Model Compression: From Theory To Practice
- LadaBERT: Lightweight Adaptation of BERT through Hybrid Model Compression
- SBNet: Sparse Blocks Network for Fast Inference
- Training Deep Neural Networks with Joint Quantization and Pruning of Weights and Activations
- Compression of Deep Convolutional Neural Networks under Joint Sparsity Constraints
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- BitNet: Bit-Regularized Deep Neural Networks
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-based Approach
- Trends and Advancements in Deep Neural Network Communication
- Detecting Dead Weights and Units in Neural Networks
- Learning Sparse & Ternary Neural Networks with Entropy-Constrained Trained Ternarization (EC2T)
- Qu-ANTI-zation: Exploiting Quantization Artifacts for Achieving Adversarial Outcomes
- A Selective Survey on Versatile Knowledge Distillation Paradigm for Neural Network Models
- Recurrent Convolution for Compact and Cost-Adjustable Neural Networks: An Empirical Study
- Information-Theoretic Understanding of Population Risk Improvement with Model Compression
- Universal Approximation Theorems of Fully Connected Binarized Neural Networks
- TOCO: A Framework for Compressing Neural Network Models Based on Tolerance Analysis
- Kernel Quantization for Efficient Network Compression
- A Targeted Acceleration and Compression Framework for Low bit Neural Networks