Convolution-Weight-Distribution Assumption: Rethinking the Criteria of Channel Pruning
arXiv:2004.11627
Abstract
Channel pruning is a popular technique for compressing convolutional neural networks (CNNs), where various pruning criteria have been proposed to remove the redundant filters. From our comprehensive experiments, we found two blind spots in the study of pruning criteria: (1) Similarity: There are some strong similarities among several primary pruning criteria that are widely cited and compared. According to these criteria, the ranks of filters'Importance Score are almost identical, resulting in similar pruned structures. (2) Applicability: The filters'Importance Score measured by some pruning criteria are too close to distinguish the network redundancy well. In this paper, we analyze these two blind spots on different types of pruning criteria with layer-wise pruning or global pruning. The analyses are based on the empirical experiments and our assumption (Convolutional Weight Distribution Assumption) that the well-trained convolutional filters each layer approximately follow a Gaussian-alike distribution. This assumption has been verified through systematic and extensive statistical tests.
Accepted by NeurIPS 2021
References in corpus (19)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- ADADELTA: An Adaptive Learning Rate Method
- Improved Regularization of Convolutional Neural Networks with Cutout
- One weird trick for parallelizing convolutional neural networks
- Rethinking the Value of Network Pruning
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- Pruning Filters for Efficient ConvNets
- Spatial Group-wise Enhance: Improving Semantic Feature Learning in Convolutional Networks
- An Entropy-based Pruning Method for CNN Compression
- Global Sparse Momentum SGD for Pruning Very Deep Neural Networks
- Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks
- AlphaGAN: Generative adversarial networks for natural image matting
- Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure
- Luck Matters: Understanding Training Dynamics of Deep ReLU Networks
- Rethinking Class-Discrimination Based CNN Channel Pruning
- COP: Customized Deep Model Compression via Regularized Correlation-Based Filter-Level Pruning
- Drop-Activation: Implicit Parameter Reduction and Harmonic Regularization
- Instance Enhancement Batch Normalization: an Adaptive Regulator of Batch Noise