Dual Precision Deep Neural Network
arXiv:2009.02191 · doi:10.1145/3430199.3430228
Abstract
On-line Precision scalability of the deep neural networks(DNNs) is a critical feature to support accuracy and complexity trade-off during the DNN inference. In this paper, we propose dual-precision DNN that includes two different precision modes in a single model, thereby supporting an on-line precision switch without re-training. The proposed two-phase training process optimizes both low- and high-precision modes.
5 pages, 4 figures, 2 tables
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep Learning with Limited Numerical Precision
- Ternary Weight Networks
- Training and Inference with Integers in Deep Neural Networks
- Trained Quantization Thresholds for Accurate and Efficient Fixed-Point Inference of Deep Neural Networks
- Post-training Quantization with Multiple Points: Mixed Precision without Mixed Precision