Sparse Networks from Scratch: Faster Training without Losing Performance
arXiv:1907.04840
Abstract
We demonstrate the possibility of what we call sparse learning: accelerated training of deep neural networks that maintain sparse weights throughout training while achieving dense performance levels. We accomplish this by developing sparse momentum, an algorithm which uses exponentially smoothed gradients (momentum) to identify layers and weights which reduce the error efficiently. Sparse momentum redistributes pruned weights across layers according to the mean momentum magnitude of each layer. Within a layer, sparse momentum grows weights according to the momentum magnitude of zero-valued weights. We demonstrate state-of-the-art sparse performance on MNIST, CIFAR-10, and ImageNet, decreasing the mean error by a relative 8%, 15%, and 6% compared to other sparse algorithms. Furthermore, we show that sparse momentum reliably reproduces dense performance levels while providing up to 5.61x faster training. In our analysis, ablations show that the benefits of momentum redistribution and growth increase with the depth and size of the network. Additionally, we find that sparse momentum is insensitive to the choice of its hyperparameters suggesting that sparse momentum is robust and easy to use.
9 page NeurIPS 2019 submission
References in corpus (14)
- Striving for Simplicity: The All Convolutional Net
- Wide Residual Networks
- To prune, or not to prune: exploring the efficacy of pruning for model compression
- SNIP: Single-shot Network Pruning based on Connection Sensitivity
- Scalable Training of Artificial Neural Networks with Adaptive Sparse Connectivity inspired by Network Science
- Generating Long Sequences with Sparse Transformers
- The State of Sparsity in Deep Neural Networks
- Bayesian Compression for Deep Learning
- Variational Dropout Sparsifies Deep Neural Networks
- Deep Rewiring: Training very sparse deep networks
- Exploring Sparsity in Recurrent Neural Networks
- Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization
- Deconstructing Lottery Tickets: Zeros, Signs, and the Supermask
- Grow and Prune Compact, Fast, and Accurate LSTMs
Cited by in corpus (69)
- On the Opportunities and Risks of Foundation Models
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Adaptive Extreme Edge Computing for Wearable Devices
- Picking Winning Tickets Before Training by Preserving Gradient Flow
- Soft Threshold Weight Reparameterization for Learnable Sparsity
- Dynamic Model Pruning with Feedback
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training and Inference
- Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch
- SpaceNet: Make Free Space For Continual Learning
- BASE Layers: Simplifying Training of Large, Sparse Models
- Discovering Neural Wirings
- Winning the Lottery with Continuous Sparsification
- Pruning Neural Networks at Initialization: Why are We Missing the Mark?
- MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the Edge
- Progressive Skeletonization: Trimming more fat from a network at initialization
- WoodFisher: Efficient Second-Order Approximation for Neural Network Compression
- The Lottery Ticket Hypothesis for Pre-trained BERT Networks
- Movement Pruning: Adaptive Sparsity by Fine-Tuning
- The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models
- Rethinking Weight Decay For Efficient Neural Network Pruning
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked Layers
- A Brain-inspired Algorithm for Training Highly Sparse Neural Networks
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse Training
- Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization
- FreezeNet: Full Performance by Reduced Storage Costs
- Selfish Sparse RNN Training
- Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N:M Transposable Masks
- Efficient Neural Network Training via Forward and Backward Propagation Sparsification
- Pruning via Iterative Ranking of Sensitivity Statistics
- Sparse Training via Boosting Pruning Plasticity with Neuroregeneration
- Towards Low-Latency Energy-Efficient Deep SNNs via Attention-Guided Compression
- Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
- Towards Learning Convolutions from Scratch
- AC/DC: Alternating Compressed/DeCompressed Training of Deep Neural Networks
- Adversarial Pruning: A Survey and Benchmark of Pruning Methods for Adversarial Robustness
- Sparse Weight Activation Training
- CAT: Compression-Aware Training for bandwidth reduction
- Improving Neural Network with Uniform Sparse Connectivity
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
- Supermasks in Superposition
- Training Deep Neural Networks with Joint Quantization and Pruning of Weights and Activations
- Powerpropagation: A sparsity inducing weight reparameterisation
- Activation function impact on Sparse Neural Networks
- Calibrate and Prune: Improving Reliability of Lottery Tickets Through Prediction Calibration
- Consistent Sparse Deep Learning: Theory and Computation
- Sparse evolutionary Deep Learning with over one million artificial neurons on commodity hardware
- Adaptive Dense-to-Sparse Paradigm for Pruning Online Recommendation System with Non-Stationary Data
- Effective Model Sparsification by Scheduled Grow-and-Prune Methods
- Know What You Don't Need: Single-Shot Meta-Pruning for Attention Heads
- Sparse Deep Learning: A New Framework Immune to Local Traps and Miscalibration
- Towards Understanding Iterative Magnitude Pruning: Why Lottery Tickets Win
- Effective Sparsification of Neural Networks with Global Sparsity Constraint
- Connectivity Matters: Neural Network Pruning Through the Lens of Effective Sparsity
- Topological Insights into Sparse Neural Networks
- Dynamic Collective Intelligence Learning: Finding Efficient Sparse Model via Refined Gradients for Pruned Weights
- A Tunable Robust Pruning Framework Through Dynamic Network Rewiring of DNNs
- Training Sparse Neural Networks using Compressed Sensing
- Towards Structured Dynamic Sparse Pre-Training of BERT
- A Bregman Learning Framework for Sparse Neural Networks
- Search Spaces for Neural Model Training
- A fast asynchronous MCMC sampler for sparse Bayesian inference
- EPIC TTS Models: Empirical Pruning Investigations Characterizing Text-To-Speech Models
- The Elastic Lottery Ticket Hypothesis
- Deconstructing the Structure of Sparse Neural Networks
- Architecture Aware Latency Constrained Sparse Neural Networks
- Efficient Sparse Artificial Neural Networks
- Neural Architecture Search via Bregman Iterations
- Learning Pruned Structure and Weights Simultaneously from Scratch: an Attention based Approach