Pruning and Quantization for Deep Neural Network Acceleration: A Survey
arXiv:2101.09671
Abstract
Deep neural networks have been applied in many applications exhibiting extraordinary abilities in the field of computer vision. However, complex network architectures challenge efficient real-time deployment and require significant computation resources and energy costs. These challenges can be overcome through optimizations such as network compression. Network compression can often be realized with little loss of accuracy. In some cases accuracy may even improve. This paper provides a survey on two types of network compression: pruning and quantization. Pruning can be categorized as static if it is performed offline or dynamic if it is performed at run-time. We compare pruning techniques and describe criteria used to remove redundant computations. We discuss trade-offs in element-wise, channel-wise, shape-wise, filter-wise, layer-wise and even network-wise pruning. Quantization reduces computations by reducing the precision of the datatype. Weights, biases, and activations may be quantized typically to 8-bit integers although lower bit width implementations are also discussed including binary neural networks. Both pruning and quantization can be used independently or combined. We compare current techniques, analyze their strengths and weaknesses, present compressed network accuracy results on a number of frameworks, and provide practical guidance for compressing networks.
References in corpus (22)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Improving neural networks by preventing co-adaptation of feature detectors
- Language Models are Few-Shot Learners
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Compressing Deep Convolutional Networks using Vector Quantization
- Trained Ternary Quantization
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- Compressing Neural Networks with the Hashing Trick
- Binary Neural Networks: A Survey
- The State of Sparsity in Deep Neural Networks
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- Training Deep Neural Networks with 8-bit Floating Point Numbers
- DSD: Dense-Sparse-Dense Training for Deep Neural Networks
- Highway and Residual Networks learn Unrolled Iterative Estimation
- Extremely Low Bit Neural Network: Squeeze the Last Bit Out with ADMM
- Dissecting the Graphcore IPU Architecture via Microbenchmarking
- Ternary Neural Networks with Fine-Grained Quantization
- Rounding Methods for Neural Networks with Low Resolution Synaptic Weights
- Training Binary Multilayer Neural Networks for Image Classification using Expectation Backpropagation
- Low-Precision Batch-Normalized Activations
- The State of Knowledge Distillation for Classification