meProp: Sparsified Back Propagation for Accelerated Deep Learning with Reduced Overfitting
arXiv:1706.06197
Abstract
We propose a simple yet effective technique for neural network learning. The forward propagation is computed as usual. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top- elements (in terms of magnitude) are kept. As a result, only rows or columns (depending on the layout) of the weight matrix are modified, leading to a linear reduction ( divided by the vector dimension) in the computational cost. Surprisingly, experimental results demonstrate that we can update only 1-4% of the weights at each back propagation pass. This does not result in a larger number of training iterations. More interestingly, the accuracy of the resulting models is actually improved rather than degraded, and a detailed analysis is given. The code is available at https://github.com/lancopku/meProp
Accepted by the 34th International Conference on Machine Learning (ICML 2017)
References in corpus (1)
Cited by in corpus (27)
- Machine Learning at the Wireless Edge: Distributed Stochastic Gradient Descent Over-the-Air
- Pruning by Explaining: A Novel Criterion for Deep Neural Network Pruning
- Local SGD Converges Fast and Communicates Little
- Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
- TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training and Inference
- Learning Sparse Networks Using Targeted Dropout
- Towards Accurate Post-Training Quantization for Vision Transformer
- ElasticTrainer: Speeding Up On-Device Training with Runtime Elastic Tensor Selection
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- SparseTrain:Leveraging Dynamic Sparsity in Training DNNs on General-Purpose SIMD Processors
- The Convergence of Sparsified Gradient Methods
- SparCML: High-Performance Sparse Communication for Machine Learning
- Sparse Weight Activation Training
- Bag-of-Words as Target for Neural Machine Translation
- Accelerating CNN Training by Pruning Activation Gradients
- Randomized Automatic Differentiation
- Toward Compact Deep Neural Networks via Energy-Aware Pruning
- Faster Neural Network Training with Approximate Tensor Operations
- Neural gradients are near-lognormal: improved quantized and sparse training
- M-FAC: Efficient Matrix-Free Approximations of Second-Order Information
- Dynamic Sparse Graph for Efficient Deep Learning
- Advancing On-Device Neural Network Training with TinyPropv2: Dynamic, Sparse, and Efficient Backpropagation
- Dynamic Collective Intelligence Learning: Finding Efficient Sparse Model via Refined Gradients for Pruned Weights
- EasyConvPooling: Random Pooling with Easy Convolution for Accelerating Training and Testing
- Accelerated CNN Training Through Gradient Approximation
- Accelerating Training using Tensor Decomposition
- Masked Training of Neural Networks with Partial Gradients