FreezeNet: Full Performance by Reduced Storage Costs
arXiv:2011.14087 · doi:10.1007/978-3-030-69544-6_41
Abstract
Pruning generates sparse networks by setting parameters to zero. In this work we improve one-shot pruning methods, applied before training, without adding any additional storage costs while preserving the sparse gradient computations. The main difference to pruning is that we do not sparsify the network's weights but learn just a few key parameters and keep the other ones fixed at their random initialized value. This mechanism is called freezing the parameters. Those frozen weights can be stored efficiently with a single 32bit random seed number. The parameters to be frozen are determined one-shot by a single for- and backward pass applied before training starts. We call the introduced method FreezeNet. In our experiments we show that FreezeNets achieve good results, especially for extreme freezing rates. Freezing weights preserves the gradient flow throughout the network and consequently, FreezeNets train better and have an increased capacity compared to their pruned counterparts. On the classification tasks MNIST and CIFAR-10/100 we outperform SNIP, in this setting the best reported one-shot pruning method, applied before training. On MNIST, FreezeNet achieves 99.2% performance of the baseline LeNet-5-Caffe architecture, while compressing the number of trained and stored parameters by a factor of x 157.
Conference Paper of the Asian Conference on Computer Vision (ACCV) 2020
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Dynamic Network Surgery for Efficient DNNs
- Global Sparse Momentum SGD for Pruning Very Deep Neural Networks
Cited by in corpus (6)
- MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the Edge
- Dimensionality Reduced Training by Pruning and Freezing Parts of a Deep Neural Network, a Survey
- An Analysis of Layer-Freezing Strategies for Enhanced Transfer Learning in YOLO Architectures
- Effective Model Sparsification by Scheduled Grow-and-Prune Methods
- COPS: Controlled Pruning Before Training Starts
- Novel Weight Update Scheme for Hardware Neural Network based on Synaptic Devices Having Abrupt LTP or LTD Characteristics