Training deep neural networks with low precision multiplications
arXiv:1412.7024
Abstract
Multipliers are the most space and power-hungry arithmetic operators of the digital implementation of deep neural networks. We train a set of state-of-the-art neural networks (Maxout networks) on three benchmark datasets: MNIST, CIFAR-10 and SVHN. They are trained with three distinct formats: floating point, fixed point and dynamic fixed point. For each of those datasets and for each of those formats, we assess the impact of the precision of the multiplications on the final error after training. We find that very low precision is sufficient not just for running trained networks but also for training them. For example, it is possible to train Maxout networks with 10 bits multiplications.
10 pages, 5 figures, Accepted as a workshop contribution at ICLR 2015
References in corpus (2)
Cited by in corpus (80)
- Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Highly Scalable Deep Learning Training System with Mixed-Precision: Training ImageNet in Four Minutes
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Hardware-oriented Approximation of Convolutional Neural Networks
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- Quantization and Deployment of Deep Neural Networks on Microcontrollers
- Origami: A 803 GOp/s/W Convolutional Network Accelerator
- A Survey on Coarse-Grained Reconfigurable Architectures from a Performance Perspective
- An Energy-Efficient FPGA-based Deconvolutional Neural Networks Accelerator for Single Image Super-Resolution
- XNOR-Net++: Improved Binary Neural Networks
- What Do Compressed Deep Neural Networks Forget?
- Scaling Deep Learning on GPU and Knights Landing clusters
- Making AI Forget You: Data Deletion in Machine Learning
- A Study of BFLOAT16 for Deep Learning Training
- NISP: Pruning Networks using Neuron Importance Score Propagation
- ShiftCNN: Generalized Low-Precision Architecture for Inference of Convolutional Neural Networks
- Pruning and Quantization for Deep Neural Network Acceleration: A Survey
- Improved training of binary networks for human pose estimation and image recognition
- Hierarchical binary CNNs for landmark localization with limited resources
- NITI: Training Integer Neural Networks Using Integer-only Arithmetic
- Quantized Convolutional Neural Networks for Mobile Devices
- DSConv: Efficient Convolution Operator
- Binary Neural Networks for Memory-Efficient and Effective Visual Place Recognition in Changing Environments
- Distributed Learning in Wireless Networks: Recent Progress and Future Challenges
- SPRING: A Sparsity-Aware Reduced-Precision Monolithic 3D CNN Accelerator Architecture for Training and Inference
- A Survey of Numerical Methods Utilizing Mixed Precision Arithmetic
- Attention-Based Neural Networks for Chroma Intra Prediction in Video Coding
- Quantizing Convolutional Neural Networks for Low-Power High-Throughput Inference Engines
- Resource-Efficient Neural Networks for Embedded Systems
- Fixed-point Quantization of Convolutional Neural Networks for Quantized Inference on Embedded Platforms
- Collaborative Execution of Deep Neural Networks on Internet of Things Devices
- Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data
- Incremental Binarization On Recurrent Neural Networks For Single-Channel Source Separation
- Quantization of Deep Neural Networks for Accumulator-constrained Processors
- VS-Quant: Per-vector Scaled Quantization for Accurate Low-Precision Neural Network Inference
- Full deep neural network training on a pruned weight budget
- Revisiting BFloat16 Training
- Organ Segmentation From Full-size CT Images Using Memory-Efficient FCN
- Term Revealing: Furthering Quantization at Run Time on Quantized DNNs
- An Ultra-Efficient Memristor-Based DNN Framework with Structured Weight Pruning and Quantization Using ADMM
- Topology of deep neural networks
- Design Flow of Accelerating Hybrid Extremely Low Bit-width Neural Network in Embedded FPGA
- Phoenix: A Low-Precision Floating-Point Quantization Oriented Architecture for Convolutional Neural Networks
- 8-bit Optimizers via Block-wise Quantization
- Matrix and tensor decompositions for training binary neural networks
- LCP: A Low-Communication Parallelization Method for Fast Neural Network Inference in Image Recognition
- Training for 'Unstable' CNN Accelerator:A Case Study on FPGA
- Visual Confusion Label Tree For Image Classification
- Computational Cost Reduction in Learned Transform Classifications
- An FPGA Accelerated Method for Training Feed-forward Neural Networks Using Alternating Direction Method of Multipliers and LSMR
- A Very Compact Embedded CNN Processor Design Based on Logarithmic Computing
- Deep Learning in Memristive Nanowire Networks
- VeriDL: Integrity Verification of Outsourced Deep Learning Services (Extended Version)
- DNN Feature Map Compression using Learned Representation over GF(2)
- Masked LARk: Masked Learning, Aggregation and Reporting worKflow
- Derivation and Analysis of Fast Bilinear Algorithms for Convolution
- How Low Can We Go: Trading Memory for Error in Low-Precision Training
- Fast Adjustable Threshold For Uniform Neural Network Quantization (Winning solution of LPIRC-II)
- Generative Design of Hardware-aware DNNs
- Efficient and Robust Machine Learning for Real-World Systems
- Numerical influence of ReLU'(0) on backpropagation
- Reducing Inference Latency with Concurrent Architectures for Image Recognition
- Integrating Deep Learning in Domain Sciences at Exascale
- Filter sharing: Efficient learning of parameters for volumetric convolutions
- Training a Binary Weight Object Detector by Knowledge Transfer for Autonomous Driving
- Learning In Practice: Reasoning About Quantization
- Deep Learning Acceleration Techniques for Real Time Mobile Vision Applications
- Conditional Neural Architecture Search
- Creating Robust Deep Neural Networks With Coded Distributed Computing for IoT Systems
- Bayesian Sparsification Methods for Deep Complex-valued Networks
- Deep Neural Network Training without Multiplications
- Low-bit Quantization of Recurrent Neural Network Language Models Using Alternating Direction Methods of Multipliers
- Robust error bounds for quantised and pruned neural networks
- Toward Compact Parameter Representations for Architecture-Agnostic Neural Network Compression
- Efficient Non-linear Calculators
- Quantization Loss Re-Learning Method
- A CNN Accelerator on FPGA Using Depthwise Separable Convolution
- The Low-Resource Double Bind: An Empirical Study of Pruning for Low-Resource Machine Translation
- Resource-Efficient Speech Mask Estimation for Multi-Channel Speech Enhancement