Learning Efficient Convolutional Networks through Network Slimming
arXiv:1708.06519
Abstract
The deployment of deep convolutional neural networks (CNNs) in many real world applications is largely hindered by their high computational cost. In this paper, we propose a novel learning scheme for CNNs to simultaneously 1) reduce the model size; 2) decrease the run-time memory footprint; and 3) lower the number of computing operations, without compromising accuracy. This is achieved by enforcing channel-level sparsity in the network in a simple but effective way. Different from many existing approaches, the proposed method directly applies to modern CNN architectures, introduces minimum overhead to the training process, and requires no special software/hardware accelerators for the resulting models. We call our approach network slimming, which takes wide and large networks as input models, but during training insignificant channels are automatically identified and pruned afterwards, yielding thin and compact models with comparable accuracy. We empirically demonstrate the effectiveness of our approach with several state-of-the-art CNN models, including VGGNet, ResNet and DenseNet, on various image classification datasets. For VGGNet, a multi-pass version of network slimming gives a 20x reduction in model size and a 5x reduction in computing operations.
Accepted by ICCV 2017
References in corpus (5)
Cited by in corpus (8)
- DBP: Discrimination Based Block-Level Pruning for Deep Model Acceleration
- Hardware-Centric AutoML for Mixed-Precision Quantization
- BLK-REW: A Unified Block-based DNN Pruning Framework using Reweighted Regularization Method
- Fractional Skipping: Towards Finer-Grained Dynamic CNN Inference
- Progressive Learning of Low-Precision Networks
- Convolutional neural networks compression with low rank and sparse tensor decompositions
- -LBI: Stochastic Split Linearized Bregman Iterations for Parsimonious Deep Learning
- Now that I can see, I can improve: Enabling data-driven finetuning of CNNs on the edge