Pruning artificial neural networks: a way to find well-generalizing, high-entropy sharp minima
arXiv:2004.14765 · doi:10.1007/978-3-030-61616-8_6
Abstract
Recently, a race towards the simplification of deep networks has begun, showing that it is effectively possible to reduce the size of these models with minimal or no performance loss. However, there is a general lack in understanding why these pruning strategies are effective. In this work, we are going to compare and analyze pruned solutions with two different pruning approaches, one-shot and gradual, showing the higher effectiveness of the latter. In particular, we find that gradual pruning allows access to narrow, well-generalizing minima, which are typically ignored when using one-shot approaches. In this work we also propose PSP-entropy, a measure to understand how a given neuron correlates to some specific learned classes. Interestingly, we observe that the features extracted by iteratively-pruned models are less correlated to specific classes, potentially making these models a better fit in transfer learning approaches.
References in corpus (5)
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- To prune, or not to prune: exploring the efficacy of pruning for model compression
- Qualitatively characterizing neural network optimization problems
- Comparing Rewinding and Fine-tuning in Neural Network Pruning
- LOss-Based SensiTivity rEgulaRization: towards deep sparse neural networks
Cited by in corpus (5)
- LOss-Based SensiTivity rEgulaRization: towards deep sparse neural networks
- SeReNe: Sensitivity based Regularization of Neurons for Structured Sparsity in Neural Networks
- A non-discriminatory approach to ethical deep learning
- The rise of the lottery heroes: why zero-shot pruning is hard
- SCoTTi: Save Computation at Training Time with an adaptive framework