Deep Learning with Limited Numerical Precision
arXiv:1502.02551
Abstract
Training of large-scale deep neural networks is often constrained by the available computational resources. We study the effect of limited precision data representation and computation on neural network training. Within the context of low-precision fixed-point computations, we observe the rounding scheme to play a crucial role in determining the network's behavior during training. Our results show that deep networks can be trained using only 16-bit wide fixed-point number representation when using stochastic rounding, and incur little to no degradation in the classification accuracy. We also demonstrate an energy-efficient hardware accelerator that implements low-precision fixed-point arithmetic with stochastic rounding.
10 pages, 6 figures, 1 table
References in corpus (2)
Cited by in corpus (36)
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Group Sparse Regularization for Deep Neural Networks
- WRPN: Wide Reduced-Precision Networks
- Compressing Recurrent Neural Network with Tensor Train
- Deep Convolutional Neural Network Inference with Floating-point Weights and Fixed-point Activations
- Scaling Deep Learning on GPU and Knights Landing clusters
- DeepIoT: Compressing Deep Neural Network Structures for Sensing Systems with a Compressor-Critic Framework
- Compressing Convolutional Neural Networks
- Overcoming Challenges in Fixed Point Training of Deep Convolutional Networks
- Optimally Scheduling CNN Convolutions for Efficient Memory Access
- Evaluating Built-in ECC of FPGA on-chip Memories for the Mitigation of Undervolting Faults
- Regularizing Activation Distribution for Training Binarized Deep Networks
- Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data
- An OpenCL(TM) Deep Learning Accelerator on Arria 10
- Low-Precision Batch-Normalized Activations
- Efficient Stochastic Inference of Bitwise Deep Neural Networks
- Mixed-precision training of deep neural networks using computational memory
- ModelHub: Towards Unified Data and Lifecycle Management for Deep Learning
- Deep Learning applied to Road Traffic Speed forecasting
- Training Bit Fully Convolutional Network for Fast Semantic Segmentation
- Understanding the Impact of Precision Quantization on the Accuracy and Energy of Neural Networks
- On the efficient representation and execution of deep acoustic models
- Tricks from Deep Learning
- Computation Error Analysis of Block Floating Point Arithmetic Oriented Convolution Neural Network Accelerator Design
- Energy-efficient Machine Learning in Silicon: A Communications-inspired Approach
- A Simple Method to Reduce Off-chip Memory Accesses on Convolutional Neural Networks
- SparCE: Sparsity aware General Purpose Core Extensions to Accelerate Deep Neural Networks
- Exploring the Imposition of Synaptic Precision Restrictions For Evolutionary Synthesis of Deep Neural Networks
- Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
- Understanding the Energy and Precision Requirements for Online Learning
- A Reconfigurable Streaming Deep Convolutional Neural Network Accelerator for Internet of Things
- CATERPILLAR: Coarse Grain Reconfigurable Architecture for Accelerating the Training of Deep Neural Networks
- NeuroTrainer: An Intelligent Memory Module for Deep Learning Training
- Structured Bayesian Compression for Deep models in mobile enabled devices for connected healthcare
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- Hardware-Software Codesign of Accurate, Multiplier-free Deep Neural Networks