Deep Rewiring: Training very sparse deep networks
arXiv:1711.05136
Abstract
Neuromorphic hardware tends to pose limits on the connectivity of deep networks that one can run on them. But also generic hardware and software implementations of deep learning run more efficiently for sparse networks. Several methods exist for pruning connections of a neural network after it was trained without connectivity constraints. We present an algorithm, DEEP R, that enables us to train directly a sparsely connected neural network. DEEP R automatically rewires the network during supervised training so that connections are there where they are most needed for the task, while its total number is all the time strictly bounded. We demonstrate that DEEP R can be used to train very sparse feedforward and recurrent neural networks on standard benchmark tasks with just a minor loss in performance. DEEP R is based on a rigorous theoretical foundation that views rewiring as stochastic sampling of network configurations from a posterior.
Accepted for publication at ICLR 2018. 10 pages (12 with references, 24 with appendix), 4 Figures in the main text. Reviews are available at: https://openreview.net/forum?id=BJ_wN01C- . This recent version contains minor corrections in the appendix
Cited by in corpus (59)
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- Sparse Networks from Scratch: Faster Training without Losing Performance
- Adaptive Extreme Edge Computing for Wearable Devices
- Picking Winning Tickets Before Training by Preserving Gradient Flow
- Stabilizing the Lottery Ticket Hypothesis
- Rigging the Lottery: Making All Tickets Winners
- Soft Threshold Weight Reparameterization for Learnable Sparsity
- Biologically inspired alternatives to backpropagation through time for learning in recurrent neural nets
- Pruning neural networks without any data by iteratively conserving synaptic flow
- Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch
- Drawing Inspiration from Biological Dendrites to Empower Artificial Neural Networks
- SpaceNet: Make Free Space For Continual Learning
- Long short-term memory and learning-to-learn in networks of spiking neurons
- Attention-based Convolutional Autoencoders for 3D-Variational Data Assimilation
- Progressive Skeletonization: Trimming more fat from a network at initialization
- Low-Memory Neural Network Training: A Technical Report
- HYDRA: Pruning Adversarially Robust Neural Networks
- The Difficulty of Training Sparse Neural Networks
- Spiking neural networks trained with backpropagation for low power neuromorphic implementation of voice activity detection
- Rethinking Weight Decay For Efficient Neural Network Pruning
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked Layers
- Looking GLAMORous: Vehicle Re-Id in Heterogeneous Cameras Networks with Global and Local Attention
- A Method for Medical Data Analysis Using the LogNNet for Clinical Decision Support Systems and Edge Computing in Healthcare
- A Brain-inspired Algorithm for Training Highly Sparse Neural Networks
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse Training
- Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization
- FreezeNet: Full Performance by Reduced Storage Costs
- Dynamic Hard Pruning of Neural Networks at the Edge of the Internet
- Selfish Sparse RNN Training
- Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N:M Transposable Masks
- Basic principles drive self-organization of brain-like connectivity structure
- Sparse Training via Boosting Pruning Plasticity with Neuroregeneration
- Pruning via Iterative Ranking of Sensitivity Statistics
- Towards Low-Latency Energy-Efficient Deep SNNs via Attention-Guided Compression
- Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
- Convergence Rates of Variational Inference in Sparse Deep Learning
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition
- AC/DC: Alternating Compressed/DeCompressed Training of Deep Neural Networks
- Improving Neural Network with Uniform Sparse Connectivity
- Campfire: Compressible, Regularization-Free, Structured Sparse Training for Hardware Accelerators
- Activation function impact on Sparse Neural Networks
- Powerpropagation: A sparsity inducing weight reparameterisation
- Intrinsically Sparse Long Short-Term Memory Networks
- Sparse evolutionary Deep Learning with over one million artificial neurons on commodity hardware
- Non-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
- Effective Model Sparsification by Scheduled Grow-and-Prune Methods
- Pruning Randomly Initialized Neural Networks with Iterative Randomization
- COPS: Controlled Pruning Before Training Starts
- Effective Sparsification of Neural Networks with Global Sparsity Constraint
- Lottery Jackpots Exist in Pre-trained Models
- Train-by-Reconnect: Decoupling Locations of Weights from their Values
- Initialization and Regularization of Factorized Neural Layers
- ESPN: Extremely Sparse Pruned Networks
- Are wider nets better given the same number of parameters?
- Simplex Closing Probabilities in Directed Graphs
- Using noise to probe recurrent neural network structure and prune synapses
- Node pruning reveals compact and optimal substructures within large networks
- Phenomenological modeling of diverse and heterogeneous synaptic dynamics at natural density
- Livewired Neural Networks: Making Neurons That Fire Together Wire Together