Compressing Deep Convolutional Networks using Vector Quantization
arXiv:1412.6115
Abstract
Deep convolutional neural networks (CNN) has become the most promising method for object recognition, repeatedly demonstrating record breaking results for image classification and object detection in recent years. However, a very deep CNN generally involves many layers with millions of parameters, making the storage of the network model to be extremely large. This prohibits the usage of deep CNNs on resource limited hardware, especially cell phones or other embedded devices. In this paper, we tackle this model storage issue by investigating information theoretical vector quantization methods for compressing the parameters of CNNs. In particular, we have found in terms of compressing the most storage demanding dense connected layers, vector quantization methods have a clear gain over existing matrix factorization methods. Simply applying k-means clustering to the weights or conducting product quantization can lead to a very good balance between model size and recognition accuracy. For the 1000-category classification task in the ImageNet challenge, we are able to achieve 16-24 times compression of the network with only 1% loss of classification accuracy using the state-of-the-art CNN.
References in corpus (3)
Cited by in corpus (62)
- FitNets: Hints for Thin Deep Nets
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Group Sparse Regularization for Deep Neural Networks
- Channel Pruning for Accelerating Very Deep Neural Networks
- Deep Adaptive Feature Embedding with Local Sample Distributions for Person Re-identification
- JALAD: Joint Accuracy- and Latency-Aware Deep Structure Decoupling for Edge-Cloud Execution
- A Comprehensive guide to Bayesian Convolutional Neural Network with Variational Inference
- An Entropy-based Pruning Method for CNN Compression
- Dynamic Network Surgery for Efficient DNNs
- ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression
- Deep Model Compression: Distilling Knowledge from Noisy Teachers
- The Power of Sparsity in Convolutional Neural Networks
- Fast YOLO: A Fast You Only Look Once System for Real-time Embedded Object Detection in Video
- DeepSZ: A Novel Framework to Compress Deep Neural Networks by Using Error-Bounded Lossy Compression
- Modeling the Resource Requirements of Convolutional Neural Networks on Mobile Devices
- Deep Learning with Low Precision by Half-wave Gaussian Quantization
- Effective Quantization Methods for Recurrent Neural Networks
- Training Skinny Deep Neural Networks with Iterative Hard Thresholding Methods
- MEC: Memory-efficient Convolution for Deep Neural Network
- ProjectionNet: Learning Efficient On-Device Deep Networks Using Neural Projections
- DeepIoT: Compressing Deep Neural Network Structures for Sensing Systems with a Compressor-Critic Framework
- Compressing Convolutional Neural Networks
- Model compression as constrained optimization, with application to neural nets. Part I: general framework
- Model compression as constrained optimization, with application to neural nets. Part II: quantization
- Collaborative Execution of Deep Neural Networks on Internet of Things Devices
- Evolutionary Synthesis of Deep Neural Networks via Synaptic Cluster-driven Genetic Encoding
- Model Distillation with Knowledge Transfer from Face Classification to Alignment and Verification
- Knowledge Projection for Deep Neural Networks
- Local Feature Detectors, Descriptors, and Image Representations: A Survey
- OpenEI: An Open Framework for Edge Intelligence
- All You Need is a Few Shifts: Designing Efficient Convolutional Neural Networks for Image Classification
- Compressing Neural Language Models by Sparse Word Representations
- Training Sparse Neural Networks
- CGaP: Continuous Growth and Pruning for Efficient Deep Learning
- Towards Compact ConvNets via Structure-Sparsity Regularized Filter Pruning
- DeepFont: Identify Your Font from An Image
- Training Bit Fully Convolutional Network for Fast Semantic Segmentation
- Sparse Neural Networks Topologies
- Efficient Semantic Scene Completion Network with Spatial Group Convolution
- Learning Student Networks via Feature Embedding
- Compressing complex convolutional neural network based on an improved deep compression algorithm
- 2-bit Model Compression of Deep Convolutional Neural Network on ASIC Engine for Image Retrieval
- DARC: Differentiable ARchitecture Compression
- Compression of Deep Neural Networks for Image Instance Retrieval
- FFT-Based Deep Learning Deployment in Embedded Systems
- Table-Based Neural Units: Fully Quantizing Networks for Multiply-Free Inference
- BasisConv: A method for compressed representation and learning in CNNs
- Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
- Efficient Inferencing of Compressed Deep Neural Networks
- CompactNet: Platform-Aware Automatic Optimization for Convolutional Neural Networks
- Lattice Rescoring Strategies for Long Short Term Memory Language Models in Speech Recognition
- Towards thinner convolutional neural networks through Gradually Global Pruning
- Information-Theoretic Understanding of Population Risk Improvement with Model Compression
- Deep Learning Acceleration Techniques for Real Time Mobile Vision Applications
- CircConv: A Structured Convolution with Low Complexity
- ADA-Tucker: Compressing Deep Neural Networks via Adaptive Dimension Adjustment Tucker Decomposition
- (Pen-) Ultimate DNN Pruning
- An Embedded Deep Learning based Word Prediction
- Seeing Convolution Through the Eyes of Finite Transformation Semigroup Theory: An Abstract Algebraic Interpretation of Convolutional Neural Networks
- MICIK: MIning Cross-Layer Inherent Similarity Knowledge for Deep Model Compression
- Scalable Compression of Deep Neural Networks
- NodeDrop: A Condition for Reducing Network Size without Effect on Output