Rigging the Lottery: Making All Tickets Winners
arXiv:1911.11134
Abstract
Many applications require sparse neural networks due to space or inference time restrictions. There is a large body of work on training dense networks to yield sparse networks for inference, but this limits the size of the largest trainable sparse model to that of the largest trainable dense model. In this paper we introduce a method to train sparse neural networks with a fixed parameter count and a fixed computational cost throughout training, without sacrificing accuracy relative to existing dense-to-sparse training methods. Our method updates the topology of the sparse network during training by using parameter magnitudes and infrequent gradient calculations. We show that this approach requires fewer floating-point operations (FLOPs) to achieve a given level of accuracy compared to prior techniques. We demonstrate state-of-the-art sparse training results on a variety of networks and datasets, including ResNet-50, MobileNets on Imagenet-2012, and RNNs on WikiText-103. Finally, we provide some insights into why allowing the topology to change during the optimization can overcome local minima encountered when the topology remains static. Code used in our work can be found in github.com/google-research/rigl.
Published in Proceedings of the 37th International Conference on Machine Learning. Code can be found in github.com/google-research/rigl
References in corpus (5)
- To prune, or not to prune: exploring the efficacy of pruning for model compression
- Pointer Sentinel Mixture Models
- The State of Sparsity in Deep Neural Networks
- Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization
- Soft Threshold Weight Reparameterization for Learnable Sparsity
Cited by in corpus (29)
- What Do Compressed Deep Neural Networks Forget?
- Pruning neural networks without any data by iteratively conserving synaptic flow
- SpaceNet: Make Free Space For Continual Learning
- Winning the Lottery with Continuous Sparsification
- Characterising Bias in Compressed Models
- Progressive Skeletonization: Trimming more fat from a network at initialization
- WoodFisher: Efficient Second-Order Approximation for Neural Network Compression
- Conservative Sparse Neural Network Embedded Frequency-Constrained Unit Commitment With Distributed Energy Resources
- Rethinking Weight Decay For Efficient Neural Network Pruning
- Robust Lottery Tickets for Pre-trained Language Models
- Network Pruning That Matters: A Case Study on Retraining Variants
- Accelerating Sparse Deep Neural Networks
- Pruning via Iterative Ranking of Sensitivity Statistics
- SparseRT: Accelerating Unstructured Sparsity on GPUs for Deep Learning Inference
- Activation function impact on Sparse Neural Networks
- Calibrate and Prune: Improving Reliability of Lottery Tickets Through Prediction Calibration
- Exploring the Impact of Model Scaling on Parameter-Efficient Tuning
- Deep Neural Network Training with Frank-Wolfe
- Topological Insights into Sparse Neural Networks
- SparseDNN: Fast Sparse Deep Learning Inference on CPUs
- GroupBERT: Enhanced Transformer Architecture with Efficient Grouped Structures
- Training Sparse Neural Networks using Compressed Sensing
- Self-Reorganizing and Rejuvenating CNNs for Increasing Model Capacity Utilization
- Livewired Neural Networks: Making Neurons That Fire Together Wire Together
- Harnessing Orthogonality to Train Low-Rank Neural Networks
- EPIC TTS Models: Empirical Pruning Investigations Characterizing Text-To-Speech Models
- Adapting the Function Approximation Architecture in Online Reinforcement Learning
- The Low-Resource Double Bind: An Empirical Study of Pruning for Low-Resource Machine Translation
- Deconstructing the Structure of Sparse Neural Networks