Rethinking floating point for deep learning
arXiv:1811.01721
Abstract
Reducing hardware overhead of neural networks for faster or lower power inference and training is an active area of research. Uniform quantization using integer multiply-add has been thoroughly investigated, which requires learning many quantization parameters, fine-tuning training or other prerequisites. Little effort is made to improve floating point relative to this baseline; it remains energy inefficient, and word size reduction yields drastic loss in needed dynamic range. We improve floating point to be more energy efficient than equivalent bit width integer hardware on a 28 nm ASIC process while retaining accuracy in 8 bits with a novel hybrid log multiply/linear add, Kulisch accumulation and tapered encodings from Gustafson's posit format. With no network retraining, and drop-in replacement of all math and float32 parameters via round-to-nearest-even only, this open-sourced 8-bit log float is within 0.9% top-1 and 0.2% top-5 accuracy of the original float32 ResNet-50 CNN model on ImageNet. Unlike int8 quantization, it is still a general purpose floating point arithmetic, interpretable out-of-the-box. Our 8/38-bit log float multiply-add is synthesized and power profiled at 28 nm at 0.96x the power and 1.12x the area of 8/32-bit integer multiply-add. In 16 bits, our log float multiply-add is 0.59x the power and 0.68x the area of IEEE 754 float16 fused multiply-add, maintaining the same signficand precision and dynamic range, proving useful for training ASICs as well.
References in corpus (3)
Cited by in corpus (18)
- Performance-Efficiency Trade-off of Low-Precision Numerical Formats in Deep Neural Networks
- PLAM: a Posit Logarithm-Approximate Multiplier
- NITI: Training Integer Neural Networks Using Integer-only Arithmetic
- PositNN: Training Deep Neural Networks with Mixed Low-Precision Posit
- A Portable Parton-Level Event Generator for the High-Luminosity LHC
- DarKnight: A Data Privacy Scheme for Training and Inference of Deep Neural Networks
- EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference
- Deep Learning Training on the Edge with Low-Precision Posits
- Real numbers, data science and chaos: How to fit any dataset with a single parameter
- Training Deep Neural Networks Using Posit Number System
- AUSN: Approximately Uniform Quantization by Adaptively Superimposing Non-uniform Distribution for Deep Neural Networks
- Shifted and Squeezed 8-bit Floating Point format for Low-Precision Training of Deep Neural Networks
- RLIBM-ALL: A Novel Polynomial Approximation Method to Produce Correctly Rounded Results for Multiple Representations and Rounding Modes
- An Efficient Deep Learning Framework for Low Rate Massive MIMO CSI Reporting
- The Efficiency Misnomer
- Towards Fully 8-bit Integer Inference for the Transformer Model
- Neural Network Training with Approximate Logarithmic Computations
- An Open-Source Framework for Efficient Numerically-Tailored Computations