Sifting out the features by pruning: Are convolutional networks the winning lottery ticket of fully connected ones?
arXiv:2104.13343
Abstract
Pruning methods can considerably reduce the size of artificial neural networks without harming their performance. In some cases, they can even uncover sub-networks that, when trained in isolation, match or surpass the test accuracy of their dense counterparts. Here we study the inductive bias that pruning imprints in such "winning lottery tickets". Focusing on visual tasks, we analyze the architecture resulting from iterative magnitude pruning of a simple fully connected network (FCN). We show that the surviving node connectivity is local in input space, and organized in patterns reminiscent of the ones found in convolutional networks (CNN). We investigate the role played by data and tasks in shaping the architecture of pruned sub-networks. Our results show that the winning lottery tickets of FCNs display the key features of CNNs. The ability of such automatic network-simplifying procedure to recover the key features "hand-crafted" in the design of CNNs suggests interesting applications to other datasets and tasks, in order to discover new and efficient architectural inductive biases.
25 pages, 18 figures; typos corrected, references added
References in corpus (14)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Language Models are Few-Shot Learners
- MLP-Mixer: An all-MLP Architecture for Vision
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- The State of Sparsity in Deep Neural Networks
- A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
- Comparing Rewinding and Fine-tuning in Neural Network Pruning
- Deep Learning for Time-Series Analysis
- Dynamic Model Pruning with Feedback
- What is the State of Neural Network Pruning?
- The Early Phase of Neural Network Training
- Towards Learning Convolutions from Scratch
- Robustness to Pruning Predicts Generalization in Deep Neural Networks