Scalable Training of Artificial Neural Networks with Adaptive Sparse Connectivity inspired by Network Science
arXiv:1707.04780 · doi:10.1038/s41467-018-04316-3
Abstract
Through the success of deep learning in various domains, artificial neural networks are currently among the most used artificial intelligence methods. Taking inspiration from the network properties of biological neural networks (e.g. sparsity, scale-freeness), we argue that (contrary to general practice) artificial neural networks, too, should not have fully-connected layers. Here we propose sparse evolutionary training of artificial neural networks, an algorithm which evolves an initial sparse topology (Erdős-Rényi random graph) of two consecutive layers of neurons into a scale-free topology, during learning. Our method replaces artificial neural networks fully-connected layers with sparse ones before training, reducing quadratically the number of parameters, with no decrease in accuracy. We demonstrate our claims on restricted Boltzmann machines, multi-layer perceptrons, and convolutional neural networks for unsupervised and supervised learning on 15 datasets. Our approach has the potential to enable artificial neural networks to scale up beyond what is currently possible.
18 pages
References in corpus (3)
Cited by in corpus (68)
- Machine Learning in Aerodynamic Shape Optimization
- Smart Anomaly Detection in Sensor Systems: A Multi-Perspective Review
- An Artificial Neuron Implemented on an Actual Quantum Processor
- Sparse Networks from Scratch: Faster Training without Losing Performance
- Adaptive Extreme Edge Computing for Wearable Devices
- Picking Winning Tickets Before Training by Preserving Gradient Flow
- Stabilizing the Lottery Ticket Hypothesis
- Rigging the Lottery: Making All Tickets Winners
- Soft Threshold Weight Reparameterization for Learnable Sparsity
- Pruning neural networks without any data by iteratively conserving synaptic flow
- Layer-adaptive sparsity for the Magnitude-based Pruning
- General Inverse Design of Thin-Film Metamaterials With Convolutional Neural Networks
- SpaceNet: Make Free Space For Continual Learning
- Deep Neural Networks using a Single Neuron: Folded-in-Time Architecture using Feedback-Modulated Delay Loops
- Sparsity through evolutionary pruning prevents neuronal networks from overfitting
- Pruning Neural Networks at Initialization: Why are We Missing the Mark?
- Progressive Skeletonization: Trimming more fat from a network at initialization
- Low-Memory Neural Network Training: A Technical Report
- The Difficulty of Training Sparse Neural Networks
- Rethinking Weight Decay For Efficient Neural Network Pruning
- Dimensionality Reduced Training by Pruning and Freezing Parts of a Deep Neural Network, a Survey
- A Brain-inspired Algorithm for Training Highly Sparse Neural Networks
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse Training
- Pre-Defined Sparse Neural Networks with Hardware Acceleration
- FreezeNet: Full Performance by Reduced Storage Costs
- Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization
- Gradient Sparsification for Efficient Wireless Federated Learning with Differential Privacy
- Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N:M Transposable Masks
- A generalization of regularized dual averaging and its dynamics
- Efficient Neural Network Training via Forward and Backward Propagation Sparsification
- Sparse Training via Boosting Pruning Plasticity with Neuroregeneration
- Structural Analysis of Sparse Neural Networks
- Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
- AC/DC: Alternating Compressed/DeCompressed Training of Deep Neural Networks
- Sparse Weight Activation Training
- Improving Neural Network with Uniform Sparse Connectivity
- Graph and Network Theory for the analysis of Criminal Networks
- Evolving Plasticity for Autonomous Learning under Changing Environmental Conditions
- Supermasks in Superposition
- Activation function impact on Sparse Neural Networks
- Sparse evolutionary Deep Learning with over one million artificial neurons on commodity hardware
- Sparse Linear Networks with a Fixed Butterfly Structure: Theory and Practice
- A supervised-learning-based strategy for optimal demand response of an HVAC System
- Pruning Randomly Initialized Neural Networks with Iterative Randomization
- Learning with Delayed Synaptic Plasticity
- Effective Model Sparsification by Scheduled Grow-and-Prune Methods
- Lottery Jackpots Exist in Pre-trained Models
- Effective Sparsification of Neural Networks with Global Sparsity Constraint
- Connectivity Matters: Neural Network Pruning Through the Lens of Effective Sparsity
- AdapMTL: Adaptive Pruning Framework for Multitask Learning Model
- Pruning deep neural networks generates a sparse, bio-inspired nonlinear controller for insect flight
- Dynamic Collective Intelligence Learning: Finding Efficient Sparse Model via Refined Gradients for Pruned Weights
- Initialization and Regularization of Factorized Neural Layers
- A Bregman Learning Framework for Sparse Neural Networks
- Leveraging Structured Pruning of Convolutional Neural Networks
- Neural Symplectic Integrator with Hamiltonian Inductive Bias for the Gravitational -body Problem
- Group-Connected Multilayer Perceptron Networks
- Pruning at Initialization -- A Sketching Perspective
- Task complexity shapes internal representations and robustness in neural networks
- Learning by Active Forgetting for Neural Networks
- The Elastic Lottery Ticket Hypothesis
- LLM-Barber: Block-Aware Rebuilder for Sparsity Mask in One-Shot for Large Language Models
- Dense for the Price of Sparse: Improved Performance of Sparsely Initialized Networks via a Subspace Offset
- Toward Building Science Discovery Machines
- Irregular Metamaterial Networks
- Towards Sobolev Pruning
- Towards Memory-Efficient Training for Extremely Large Output Spaces -- Learning with 500k Labels on a Single Commodity GPU
- Optimizing Connectivity through Network Gradients for Restricted Boltzmann Machines