One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
arXiv:1906.02773
Abstract
The success of lottery ticket initializations (Frankle and Carbin, 2019) suggests that small, sparsified networks can be trained so long as the network is initialized appropriately. Unfortunately, finding these "winning ticket" initializations is computationally expensive. One potential solution is to reuse the same winning tickets across a variety of datasets and optimizers. However, the generality of winning ticket initializations remains unclear. Here, we attempt to answer this question by generating winning tickets for one training configuration (optimizer and dataset) and evaluating their performance on another configuration. Perhaps surprisingly, we found that, within the natural images domain, winning ticket initializations generalized across a variety of datasets, including Fashion MNIST, SVHN, CIFAR-10/100, ImageNet, and Places365, often achieving performance close to that of winning tickets generated on the same dataset. Moreover, winning tickets generated using larger datasets consistently transferred better than those generated using smaller datasets. We also found that winning ticket initializations generalize across optimizers with high performance. These results suggest that winning ticket initializations generated by sufficiently large datasets contain inductive biases generic to neural networks more broadly which improve training across many settings and provide hope for the development of better initialization methods.
NeurIPS 2019
References in corpus (10)
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- How transferable are features in deep neural networks?
- Rethinking the Value of Network Pruning
- A Convergence Theory for Deep Learning via Over-Parameterization
- Pruning Filters for Efficient ConvNets
- To prune, or not to prune: exploring the efficacy of pruning for model compression
- The State of Sparsity in Deep Neural Networks
- Variational Dropout Sparsifies Deep Neural Networks
- On the Power of Over-parametrization in Neural Networks with Quadratic Activation
- Critical initialisation for deep signal propagation in noisy rectifier neural networks
Cited by in corpus (27)
- TabTransformer: Tabular Data Modeling Using Contextual Embeddings
- Optimization for deep learning: theory and algorithms
- Drawing Early-Bird Tickets: Towards More Efficient Training of Deep Networks
- When BERT Plays the Lottery, All Tickets Are Winning
- Policy Manifold Search: Exploring the Manifold Hypothesis for Diversity-based Neuroevolution
- Rethinking Weight Decay For Efficient Neural Network Pruning
- Learning credit assignment
- Achieving Adversarial Robustness via Sparsity
- An Experimental Study of the Impact of Pre-training on the Pruning of a Convolutional Neural Network
- Pruning via Iterative Ranking of Sensitivity Statistics
- Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
- Exploring Lottery Ticket Hypothesis in Media Recommender Systems
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
- The Differentially Private Lottery Ticket Mechanism
- On Iterative Neural Network Pruning, Reinitialization, and the Similarity of Masks
- Finding trainable sparse networks through Neural Tangent Transfer
- Identifying Critical Neurons in ANN Architectures using Mixed Integer Programming
- Bespoke vs. Prêt-à-Porter Lottery Tickets: Exploiting Mask Similarity for Trainable Sub-Network Finding
- Super Tickets in Pre-Trained Language Models: From Model Compression to Improving Generalization
- Juvenile state hypothesis: What we can learn from lottery ticket hypothesis researches?
- The curious case of developmental BERTology: On sparsity, transfer learning, generalization and the brain
- Ultra-light deep MIR by trimming lottery tickets
- Data-dependent Pruning to find the Winning Lottery Ticket
- Sifting out the features by pruning: Are convolutional networks the winning lottery ticket of fully connected ones?
- Membership Inference Attacks on Lottery Ticket Networks
- Joint Learning of Neural Transfer and Architecture Adaptation for Image Recognition
- Out-of-the-box channel pruned networks