Learning Sparse Neural Networks through Regularization
arXiv:1712.01312
Abstract
We propose a practical method for norm regularization for neural networks: pruning the network during training by encouraging weights to become exactly zero. Such regularization is interesting since (1) it can greatly speed up training and inference, and (2) it can improve generalization. AIC and BIC, well-known model selection criteria, are special cases of regularization. However, since the norm of weights is non-differentiable, we cannot incorporate it directly as a regularization term in the objective function. We propose a solution through the inclusion of a collection of non-negative stochastic gates, which collectively determine which weights to set to zero. We show that, somewhat surprisingly, for certain distributions over the gates, the expected norm of the resulting gated weights is differentiable with respect to the distribution parameters. We further propose the \emph{hard concrete} distribution for the gates, which is obtained by "stretching" a binary concrete distribution and then transforming its samples with a hard-sigmoid. The parameters of the distribution over the gates can then be jointly optimized with the original network parameters. As a result our method allows for straightforward and efficient learning of model structures with stochastic gradient descent and allows for conditional computation in a principled way. We perform various experiments to demonstrate the effectiveness of the resulting approach and regularizer.
Published as a conference paper at the International Conference on Learning Representations (ICLR) 2018
References in corpus (3)
Cited by in corpus (50)
- Ablation Studies in Artificial Neural Networks
- Training with Quantization Noise for Extreme Model Compression
- Maximizing information from chemical engineering data sets: Applications to machine learning
- What Do Compressed Deep Neural Networks Forget?
- A Low Effort Approach to Structured CNN Design Using PCA
- Learning Sparse Networks Using Targeted Dropout
- SEALion: a Framework for Neural Network Inference on Encrypted Data
- Rethinking Weight Decay For Efficient Neural Network Pruning
- How fine can fine-tuning be? Learning efficient language models
- A Brain-inspired Algorithm for Training Highly Sparse Neural Networks
- Selfish Sparse RNN Training
- Learning to Drop: Robust Graph Neural Network via Topological Denoising
- Pruning-aware Sparse Regularization for Network Pruning
- Fast Graph Attention Networks Using Effective Resistance Based Graph Sparsification
- DeepHoyer: Learning Sparser Neural Network with Differentiable Scale-Invariant Sparsity Measures
- Full deep neural network training on a pruned weight budget
- Rethinking Class-Discrimination Based CNN Channel Pruning
- GASL: Guided Attention for Sparsity Learning in Deep Neural Networks
- HEMP: High-order Entropy Minimization for neural network comPression
- SparseRT: Accelerating Unstructured Sparsity on GPUs for Deep Learning Inference
- Powerpropagation: A sparsity inducing weight reparameterisation
- Sparseout: Controlling Sparsity in Deep Networks
- Revisiting Loss Modelling for Unstructured Pruning
- Learning Sparse Neural Networks via Sensitivity-Driven Regularization
- Artificial neural networks condensation: A strategy to facilitate adaption of machine learning in medical settings by reducing computational burden
- A Winning Hand: Compressing Deep Networks Can Improve Out-Of-Distribution Robustness
- Pruning deep neural networks generates a sparse, bio-inspired nonlinear controller for insect flight
- Fine-Grained Stochastic Architecture Search
- Training Neural Networks with Fixed Sparse Masks
- SparseDNN: Fast Sparse Deep Learning Inference on CPUs
- Learning Sparse & Ternary Neural Networks with Entropy-Constrained Trained Ternarization (EC2T)
- Principal Component Networks: Parameter Reduction Early in Training
- Searching to Sparsify Tensor Decomposition for N-ary Relational Data
- Block-wise Dynamic Sparseness
- Population-based Gradient Descent Weight Learning for Graph Coloring Problems
- Stochastic In-Face Frank-Wolfe Methods for Non-Convex Optimization and Sparse Neural Network Training
- BWCP: Probabilistic Learning-to-Prune Channels for ConvNets via Batch Whitening
- A Bregman Learning Framework for Sparse Neural Networks
- Neuron ranking -- an informed way to condense convolutional neural networks architecture
- The Role of Regularization in Shaping Weight and Node Pruning Dependency and Dynamics
- Sparsifying networks by traversing Geodesics
- How Does BN Increase Collapsed Neural Network Filters?
- Soft Attention: Does it Actually Help to Learn Social Interactions in Pedestrian Trajectory Prediction?
- -based Sparse Canonical Correlation Analysis
- A Note on Latency Variability of Deep Neural Networks for Mobile Inference
- HALO: Learning to Prune Neural Networks with Shrinkage
- Variational Dropout Sparsification for Particle Identification speed-up
- Which is Making the Contribution: Modulating Unimodal and Cross-modal Dynamics for Multimodal Sentiment Analysis
- DAC: Data-free Automatic Acceleration of Convolutional Networks
- The Low-Resource Double Bind: An Empirical Study of Pruning for Low-Resource Machine Translation