Gator: Customizable Channel Pruning of Neural Networks with Gating
arXiv:2205.15404 · doi:10.1007/978-3-030-86380-7_5
Abstract
The rise of neural network (NN) applications has prompted an increased interest in compression, with a particular focus on channel pruning, which does not require any additional hardware. Most pruning methods employ either single-layer operations or global schemes to determine which channels to remove followed by fine-tuning of the network. In this paper we present Gator, a channel-pruning method which temporarily adds learned gating mechanisms for pruning of individual channels, and which is trained with an additional auxiliary loss, aimed at reducing the computational cost due to memory, (theoretical) speedup (in terms of FLOPs), and practical, hardware-specific speedup. Gator introduces a new formulation of dependencies between NN layers which, in contrast to most previous methods, enables pruning of non-sequential parts, such as layers on ResNet's highway, and even removing entire ResNet blocks. Gator's pruning for ResNet-50 trained on ImageNet produces state-of-the-art (SOTA) results, such as 50% FLOPs reduction with only 0.4%-drop in top-5 accuracy. Also, Gator outperforms previous pruning models, in terms of GPU latency by running 1.4 times faster. Furthermore, Gator achieves improved top-5 accuracy results, compared to MobileNetV2 and SqueezeNet, for similar runtimes. The source code of this work is available at: https://github.com/EliPassov/gator.
14 pages, 3 figures. The version that appeared in ICANN is an earlier version
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Neural Architecture Search with Reinforcement Learning
- Categorical Reparameterization with Gumbel-Softmax
- Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
- AMC: AutoML for Model Compression and Acceleration on Mobile Devices
- Pruning Filters for Efficient ConvNets
- Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights
- Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition
- Discrimination-aware Channel Pruning for Deep Neural Networks
- Channel Gating Neural Networks
- What is the State of Neural Network Pruning?