Plant 'n' Seek: Can You Find the Winning Ticket?
arXiv:2111.11153
Abstract
The lottery ticket hypothesis has sparked the rapid development of pruning algorithms that aim to reduce the computational costs associated with deep learning during training and model deployment. Currently, such algorithms are primarily evaluated on imaging data, for which we lack ground truth information and thus the understanding of how sparse lottery tickets could be. To fill this gap, we develop a framework that allows us to plant and hide winning tickets with desirable properties in randomly initialized neural networks. To analyze the ability of state-of-the-art pruning to identify tickets of extreme sparsity, we design and hide such tickets solving four challenging tasks. In extensive experiments, we observe similar trends as in imaging studies, indicating that our framework can provide transferable insights into realistic problems. Additionally, we can now see beyond such relative trends and highlight limitations of current pruning methods. Based on our results, we conclude that the current limitations in ticket sparsity are likely of algorithmic rather than fundamental nature. We anticipate that comparisons to planted tickets will facilitate future developments of efficient pruning algorithms.
References in corpus (12)
- Reconciling modern machine learning practice and the bias-variance trade-off
- SNIP: Single-shot Network Pruning based on Connection Sensitivity
- Comparing Rewinding and Fine-tuning in Neural Network Pruning
- Picking Winning Tickets Before Training by Preserving Gradient Flow
- Optimal approximation of continuous functions by very deep ReLU networks
- What Can Neural Networks Reason About?
- Proving the Lottery Ticket Hypothesis: Pruning is All You Need
- Winning the Lottery with Continuous Sparsification
- Pruning via Iterative Ranking of Sensitivity Statistics
- Multi-Prize Lottery Ticket Hypothesis: Finding Accurate Binary Neural Networks by Pruning A Randomly Weighted Network
- Lottery Ticket Preserves Weight Correlation: Is It Desirable or Not?
- Generalized Dropout