CPT: Efficient Deep Neural Network Training via Cyclic Precision
arXiv:2101.09868
Abstract
Low-precision deep neural network (DNN) training has gained tremendous attention as reducing precision is one of the most effective knobs for boosting DNNs' training time/energy efficiency. In this paper, we attempt to explore low-precision training from a new perspective as inspired by recent findings in understanding DNN training: we conjecture that DNNs' precision might have a similar effect as the learning rate during DNN training, and advocate dynamic precision along the training trajectory for further boosting the time/energy efficiency of DNN training. Specifically, we propose Cyclic Precision Training (CPT) to cyclically vary the precision between two boundary values which can be identified using a simple precision range test within the first few training epochs. Extensive simulations and ablation studies on five datasets and eleven models demonstrate that CPT's effectiveness is consistent across various models/tasks (including classification and language modeling). Furthermore, through experiments and visualization we show that CPT helps to (1) converge to a wider minima with a lower generalization error and (2) reduce training variance which we believe opens up a new design knob for simultaneously improving the optimization and efficiency of DNN training. Our codes are available at: https://github.com/RICE-EIC/CPT.
Accepted at ICLR 2021 (Spotlight)
References in corpus (21)
- Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Mixed Precision Training
- Ternary Weight Networks
- Trained Ternary Quantization
- TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning
- Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks
- Pointer Sentinel Mixture Models
- Regularizing and Optimizing LSTM Language Models
- Learned Step Size Quantization
- Adding Gradient Noise Improves Learning for Very Deep Networks
- Training Deep Neural Networks with 8-bit Floating Point Numbers
- Scalable Methods for 8-bit Training of Neural Networks
- WRPN: Wide Reduced-Precision Networks
- Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy
- On the Spectral Bias of Neural Networks
- Deep -Means: Re-Training and Parameter Sharing with Harder Cluster Assignments for Compressing Deep Convolutions
- FracTrain: Fractionally Squeezing Bit Savings Both Temporally and Spatially for Efficient DNN Training
- Fractional Skipping: Towards Finer-Grained Dynamic CNN Inference
- DNQ: Dynamic Network Quantization