Compression of Deep Convolutional Neural Networks for Fast and Low Power Mobile Applications
arXiv:1511.06530
Abstract
Although the latest high-end smartphone has powerful CPU and GPU, running deeper convolutional neural networks (CNNs) for complex tasks such as ImageNet classification on mobile devices is challenging. To deploy deep CNNs on mobile devices, we present a simple and effective scheme to compress the entire CNN, which we call one-shot whole network compression. The proposed scheme consists of three steps: (1) rank selection with variational Bayesian matrix factorization, (2) Tucker decomposition on kernel tensor, and (3) fine-tuning to recover accumulated loss of accuracy, and each step can be easily implemented using publicly available tools. We demonstrate the effectiveness of the proposed scheme by testing the performance of various compressed CNNs (AlexNet, VGGS, GoogLeNet, and VGG-16) on the smartphone. Significant reductions in model size, runtime, and energy consumption are obtained, at the cost of small loss in accuracy. In addition, we address the important implementation level issue on 1?1 convolution, which is a key operation of inception module of GoogLeNet as well as CNNs compressed by our proposed scheme.
References in corpus (5)
- Improving neural networks by preventing co-adaptation of feature detectors
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Compressing Deep Convolutional Networks using Vector Quantization
- Compressing Neural Networks with the Hashing Trick
- Speeding up Convolutional Neural Networks with Low Rank Expansions
Cited by in corpus (28)
- Pruning Convolutional Neural Networks for Resource Efficient Inference
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Interleaved Group Convolutions for Deep Neural Networks
- Towards the Limit of Network Quantization
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- An Experimental Study of Reduced-Voltage Operation in Modern FPGAs for Neural Network Acceleration
- MEC: Memory-efficient Convolution for Deep Neural Network
- Hybrid Tensor Decomposition in Neural Network Compression
- Edge AI: On-Demand Accelerating Deep Neural Network Inference via Edge Computing
- Learned Threshold Pruning
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- The ZipML Framework for Training Models with End-to-End Low Precision: The Cans, the Cannots, and a Little Bit of Deep Learning
- Exploiting Errors for Efficiency: A Survey from Circuits to Algorithms
- Dynamic Zoom-in Network for Fast Object Detection in Large Images
- Coordinating Filters for Faster Deep Neural Networks
- Tensor Contraction Layers for Parsimonious Deep Nets
- CUP: Cluster Pruning for Compressing Deep Neural Networks
- PENNI: Pruned Kernel Sharing for Efficient CNN Inference
- Depth-wise Decomposition for Accelerating Separable Convolutions in Efficient Convolutional Neural Networks
- Getting deep recommenders fit: Bloom embeddings for sparse binary input/output networks
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- Compressing CNN Kernels for Videos Using Tucker Decompositions: Towards Lightweight CNN Applications
- Doubly Nested Network for Resource-Efficient Inference
- TDC: Towards Extremely Efficient CNNs on GPUs via Hardware-Aware Tucker Decomposition
- RRNet: Repetition-Reduction Network for Energy Efficient Decoder of Depth Estimation
- A Low-Compexity Deep Learning Framework For Acoustic Scene Classification
- Auto Deep Compression by Reinforcement Learning Based Actor-Critic Structure
- Tensor-based framework for training flexible neural networks