Flexpoint: An Adaptive Numerical Format for Efficient Training of Deep Neural Networks
arXiv:1711.02213
Abstract
Deep neural networks are commonly developed and trained in 32-bit floating point format. Significant gains in performance and energy efficiency could be realized by training and inference in numerical formats optimized for deep learning. Despite advances in limited precision inference in recent years, training of neural networks in low bit-width remains a challenging problem. Here we present the Flexpoint data format, aiming at a complete replacement of 32-bit floating point format training and inference, designed to support modern deep network topologies without modifications. Flexpoint tensors have a shared exponent that is dynamically adjusted to minimize overflows and maximize available dynamic range. We validate Flexpoint by training AlexNet, a deep residual network and a generative adversarial network, using a simulator implemented with the neon deep learning framework. We demonstrate that 16-bit Flexpoint closely matches 32-bit floating point in training all three models, without any need for tuning of model hyperparameters. Our results suggest Flexpoint as a promising numerical format for future hardware for training and inference.
14 pages, 5 figures, accepted in Neural Information Processing Systems 2017
References in corpus (7)
- Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Deep Learning with Limited Numerical Precision
- Convolutional Neural Networks using Logarithmic Data Representation
- Mixed Low-precision Deep Learning Inference using Dynamic Fixed Point
- Accelerating Deep Convolutional Networks using low-precision and sparsity
Cited by in corpus (14)
- Zero-Shot Text-to-Image Generation
- NVIDIA Tensor Core Programmability, Performance & Precision
- TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training and Inference
- Accelerating Generalized Linear Models with MLWeaving: A One-Size-Fits-All System for Any-precision Learning (Technical Report)
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Quantization of Deep Neural Networks for Accumulator-constrained Processors
- Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
- Hierarchical Training of Deep Neural Networks Using Early Exiting
- DaCapo: Accelerating Continuous Learning in Autonomous Systems for Video Analytics
- Elastic Significant Bit Quantization and Acceleration for Deep Neural Networks
- MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators
- TaxoNN: A Light-Weight Accelerator for Deep Neural Network Training
- Approximations in Deep Learning
- Design of Reconfigurable Multi-Operand Adder for Massively Parallel Processing