A Brain-inspired Algorithm for Training Highly Sparse Neural Networks
arXiv:1903.07138 · doi:10.1007/s10994-022-06266-w
Abstract
Sparse neural networks attract increasing interest as they exhibit comparable performance to their dense counterparts while being computationally efficient. Pruning the dense neural networks is among the most widely used methods to obtain a sparse neural network. Driven by the high training cost of such methods that can be unaffordable for a low-resource device, training sparse neural networks sparsely from scratch has recently gained attention. However, existing sparse training algorithms suffer from various issues, including poor performance in high sparsity scenarios, computing dense gradient information during training, or pure random topology search. In this paper, inspired by the evolution of the biological brain and the Hebbian learning theory, we present a new sparse training approach that evolves sparse neural networks according to the behavior of neurons in the network. Concretely, by exploiting the cosine similarity metric to measure the importance of the connections, our proposed method, Cosine similarity-based and Random Topology Exploration (CTRE), evolves the topology of sparse neural networks by adding the most important connections to the network without calculating dense gradient in the backward. We carried out different experiments on eight datasets, including tabular, image, and text datasets, and demonstrate that our proposed method outperforms several state-of-the-art sparse training algorithms in extremely sparse neural networks by a large gap. The implementation code is available on https://github.com/zahraatashgahi/CTRE
References in corpus (21)
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- SNIP: Single-shot Network Pruning based on Connection Sensitivity
- The State of Sparsity in Deep Neural Networks
- Deep Learning Scaling is Predictable, Empirically
- Sparse Networks from Scratch: Faster Training without Losing Performance
- Deep Rewiring: Training very sparse deep networks
- Learning Sparse Neural Networks through Regularization
- Picking Winning Tickets Before Training by Preserving Gradient Flow
- Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization
- Assessing the Scalability of Biologically-Motivated Deep Learning Algorithms and Architectures
- Soft Threshold Weight Reparameterization for Learnable Sparsity
- Pruning neural networks without any data by iteratively conserving synaptic flow
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular Data
- Winning the Lottery with Continuous Sparsification
- Progressive Skeletonization: Trimming more fat from a network at initialization
- RadiX-Net: Structured Sparse Matrices for Deep Neural Networks
- EigenDamage: Structured Pruning in the Kronecker-Factored Eigenbasis
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked Layers
- Top-KAST: Top-K Always Sparse Training
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse Training
- Sparse Weight Activation Training