Differentiable Sparsification for Deep Neural Networks
arXiv:1910.03201
Abstract
Deep neural networks have significantly alleviated the burden of feature engineering, but comparable efforts are now required to determine effective architectures for these networks. Furthermore, as network sizes have become excessively large, a substantial amount of resources is invested in reducing their sizes. These challenges can be effectively addressed through the sparsification of over-complete models. In this study, we propose a fully differentiable sparsification method for deep neural networks, which can zero out unimportant parameters by directly optimizing a regularized objective function with stochastic gradient descent. Consequently, the proposed method can learn both the sparsified structure and weights of a network in an end-to-end manner. It can be directly applied to various modern deep neural networks and requires minimal modification to the training process. To the best of our knowledge, this is the first fully differentiable sparsification method.
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Neural Architecture Search with Reinforcement Learning
- DARTS: Differentiable Architecture Search
- Rethinking the Value of Network Pruning
- Learning Structured Sparsity in Deep Neural Networks
- Highway and Residual Networks learn Unrolled Iterative Estimation
- Discovering Neural Wirings
- Connectivity Learning in Multi-Branch Networks
Cited by in corpus (4)
- Soft Threshold Weight Reparameterization for Learnable Sparsity
- Towards Improving the Consistency, Efficiency, and Flexibility of Differentiable Neural Architecture Search
- Convolutional neural networks compression with low rank and sparse tensor decompositions
- Embedding Differentiable Sparsity into Deep Neural Network