Quantized Convolutional Neural Networks for Mobile Devices
arXiv:1512.06473
Abstract
Recently, convolutional neural networks (CNN) have demonstrated impressive performance in various computer vision tasks. However, high performance hardware is typically indispensable for the application of CNN models due to the high computation complexity, which prohibits their further extensions. In this paper, we propose an efficient framework, namely Quantized CNN, to simultaneously speed-up the computation and reduce the storage and memory overhead of CNN models. Both filter kernels in convolutional layers and weighting matrices in fully-connected layers are quantized, aiming at minimizing the estimation error of each layer's response. Extensive experiments on the ILSVRC-12 benchmark demonstrate 4~6x speed-up and 15~20x compression with merely one percentage loss of classification accuracy. With our quantized CNN model, even mobile devices can accurately classify images within one second.
Accepted by the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2016
References in corpus (7)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Going Deeper with Convolutions
- Compressing Deep Convolutional Networks using Vector Quantization
- Compressing Neural Networks with the Hashing Trick
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Compressing Convolutional Neural Networks
Cited by in corpus (23)
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Model compression via distillation and quantization
- An Entropy-based Pruning Method for CNN Compression
- ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression
- Mobile Video Object Detection with Temporally-Aware Feature Maps
- BlockDrop: Dynamic Inference Paths in Residual Networks
- AdaptivFloat: A Floating-point based Data Type for Resilient Deep Learning Inference
- Binary Ensemble Neural Network: More Bits per Network or More Networks per Bit?
- Full deep neural network training on a pruned weight budget
- On the efficient representation and execution of deep acoustic models
- Data Efficient Stagewise Knowledge Distillation
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- Manifold Regularized Dynamic Network Pruning
- GroupReduce: Block-Wise Low-Rank Approximation for Neural Language Model Shrinking
- Pre-Quantized Deep Learning Models Codified in ONNX to Enable Hardware/Software Co-Design
- Deep Neural Network Approximation using Tensor Sketching
- Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
- Convolutional neural networks compression with low rank and sparse tensor decompositions
- Compact retail shelf segmentation for mobile deployment
- MoRS: An Approximate Fault Modelling Framework for Reduced-Voltage SRAMs
- A Novel ANN Structure for Image Recognition
- Accuracy to Throughput Trade-offs for Reduced Precision Neural Networks on Reconfigurable Logic