Pruning Ternary Quantization
arXiv:2107.10998
Abstract
Inference time, model size, and accuracy are three key factors in deep model compression. Most of the existing work addresses these three key factors separately as it is difficult to optimize them all at the same time. For example, low-bit quantization aims at obtaining a faster model; weight sharing quantization aims at improving compression ratio and accuracy; and mixed-precision quantization aims at balancing accuracy and inference time. To simultaneously optimize bit-width, model size, and accuracy, we propose pruning ternary quantization (PTQ): a simple, effective, symmetric ternary quantization method. We integrate L2 normalization, pruning, and the weight decay term to reduce the weight discrepancy in the gradient estimator during quantization, thus producing highly compressed ternary weights. Our method brings the highest test accuracy and the highest compression ratio. For example, it produces a 939kb (49) 2bit ternary ResNet-18 model with only 4\% accuracy drop on the ImageNet dataset. It compresses 170MB Mask R-CNN to 5MB (34) with only 2.8\% average precision drop. Our method is verified on image classification, object detection/segmentation tasks with different network structures such as ResNet-18, ResNet-50, and MobileNetV2.
Merged with Hyperspherical Quantization: Toward Smaller and More Accurate Models (arXiv:2212.12653.)
References in corpus (21)
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Binarized Neural Networks
- Ternary Weight Networks
- Trained Ternary Quantization
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- PACT: Parameterized Clipping Activation for Quantized Neural Networks
- Pruning Filters for Efficient ConvNets
- Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights
- Convolutional Neural Networks using Logarithmic Data Representation
- Learned Step Size Quantization
- Compression-aware Training of Deep Networks
- Orthogonal Weight Normalization: Solution to Optimization over Multiple Dependent Stiefel Manifolds in Deep Neural Networks
- Regularizing CNNs with Locally Constrained Decorrelations
- Extremely Low Bit Neural Network: Squeeze the Last Bit Out with ADMM
- ProxQuant: Quantized Neural Networks via Proximal Operators
- Generalized BackPropagation, Étude De Cas: Orthogonality
- A Unified Framework of DNN Weight Pruning and Weight Clustering/Quantization Using ADMM
- Optimization on Submanifolds of Convolution Kernels in CNNs
- Layer-compensated Pruning for Resource-constrained Convolutional Neural Networks
- Bridging the Accuracy Gap for 2-bit Quantized Neural Networks (QNN)