Mixed Precision Training
arXiv:1710.03740
Abstract
Deep neural networks have enabled progress in a wide variety of applications. Growing the size of the neural network typically results in improved accuracy. As model sizes grow, the memory and compute requirements for training these models also increases. We introduce a technique to train deep neural networks using half precision floating point numbers. In our technique, weights, activations and gradients are stored in IEEE half-precision format. Half-precision floating numbers have limited numerical range compared to single-precision numbers. We propose two techniques to handle this loss of information. Firstly, we recommend maintaining a single-precision copy of the weights that accumulates the gradients after each optimizer step. This single-precision copy is rounded to half-precision format during training. Secondly, we propose scaling the loss appropriately to handle the loss of information with half-precision gradients. We demonstrate that this approach works for a wide variety of models including convolution neural networks, recurrent neural networks and generative adversarial networks. This technique works for large scale models with more than 100 million parameters trained on large datasets. Using this approach, we can reduce the memory consumption of deep learning models by nearly 2x. In future processors, we can also expect a significant computation speedup using half-precision hardware units.
Published as a conference paper at ICLR 2018
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Deep Learning with Limited Numerical Precision
- WRPN: Wide Reduced-Precision Networks
- Effective Quantization Methods for Recurrent Neural Networks
- Neural Speech Recognizer: Acoustic-to-Word LSTM Model for Large Vocabulary Speech Recognition
- Recurrent Neural Networks With Limited Numerical Precision
Cited by in corpus (273)
- Learning Transferable Visual Models From Natural Language Supervision
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- On the Opportunities and Risks of Foundation Models
- Diffusion Models Beat GANs on Image Synthesis
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Zero-Shot Text-to-Image Generation
- fastai: A Layered API for Deep Learning
- Linformer: Self-Attention with Linear Complexity
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- PadChest: A large chest x-ray image dataset with multi-label annotated reports
- TorchIO: A Python library for efficient loading, preprocessing, augmentation and patch-based sampling of medical images in deep learning
- Fast is better than free: Revisiting adversarial training
- Generating Long Sequences with Sparse Transformers
- Highly Scalable Deep Learning Training System with Mixed-Precision: Training ImageNet in Four Minutes
- Slalom: Fast, Verifiable and Private Execution of Neural Networks in Trusted Hardware
- Deep Residual Learning in Spiking Neural Networks
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- Quantization and Deployment of Deep Neural Networks on Microcontrollers
- FILIP: Fine-grained Interactive Language-Image Pre-Training
- DMT: Dynamic Mutual Training for Semi-Supervised Learning
- Unifying Vision-and-Language Tasks via Text Generation
- NeMo: a toolkit for building AI applications using Neural Modules
- CornerNet-Lite: Efficient Keypoint Based Object Detection
- Survey of Machine Learning Accelerators
- Neural Program Repair with Execution-based Backpropagation
- Bag of Tricks for Image Classification with Convolutional Neural Networks
- Pyramidal Convolution: Rethinking Convolutional Neural Networks for Visual Recognition
- A Practical Survey on Faster and Lighter Transformers
- CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
- Mixed Precision Training of Convolutional Neural Networks using Integer Operations
- Mixed-precision deep learning based on computational memory
- Generalized Radiograph Representation Learning via Cross-supervision between Images and Free-text Radiology Reports
- Classification of Brain Tumours in MR Images using Deep Spatiospatial Models
- Scaling Distributed Machine Learning with In-Network Aggregation
- GraphVite: A High-Performance CPU-GPU Hybrid System for Node Embedding
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
- Selection via Proxy: Efficient Data Selection for Deep Learning
- fairseq S2T: Fast Speech-to-Text Modeling with fairseq
- What Do Compressed Deep Neural Networks Forget?
- NEZHA: Neural Contextualized Representation for Chinese Language Understanding
- 3D RoI-aware U-Net for Accurate and Efficient Colorectal Tumor Segmentation
- Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks
- TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training and Inference
- Autoencoders on FPGAs for real-time, unsupervised new physics detection at 40 MHz at the Large Hadron Collider
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
- Massively Distributed SGD: ImageNet/ResNet-50 Training in a Flash
- High-Accuracy Low-Precision Training
- Rethinking Positional Encoding in Language Pre-training
- Non-Autoregressive Machine Translation with Disentangled Context Transformer
- A Study of BFLOAT16 for Deep Learning Training
- ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning
- Self-attending RNN for Speech Enhancement to Improve Cross-corpus Generalization
- SNIPER: Efficient Multi-Scale Training
- Dissecting Tensor Cores via Microbenchmarks: Latency, Throughput and Numeric Behaviors
- Charged particle tracking via edge-classifying interaction networks
- EmoNeXt: an Adapted ConvNeXt for Facial Emotion Recognition
- Pruning and Quantization for Deep Neural Network Acceleration: A Survey
- Optimizing Network Performance for Distributed DNN Training on GPU Clusters: ImageNet/AlexNet Training in 1.5 Minutes
- On-Device Machine Learning: An Algorithms and Learning Theory Perspective
- Dreaming to Distill: Data-free Knowledge Transfer via DeepInversion
- Towards an astronomical foundation model for stars with a Transformer-based model
- Understanding and Improving Fast Adversarial Training
- Jasper: An End-to-End Convolutional Neural Acoustic Model
- ReLeQ: A Reinforcement Learning Approach for Deep Quantization of Neural Networks
- Mixed-Precision Training for NLP and Speech Recognition with OpenSeq2Seq
- Mixed Precision Training With 8-bit Floating Point
- Achieving Peak Performance for Large Language Models: A Systematic Review
- NITI: Training Integer Neural Networks Using Integer-only Arithmetic
- GaNDLF: A Generally Nuanced Deep Learning Framework for Scalable End-to-End Clinical Workflows in Medical Imaging
- Compounding the Performance Improvements of Assembled Techniques in a Convolutional Neural Network
- Low-Memory Neural Network Training: A Technical Report
- Scale out for large minibatch SGD: Residual network training on ImageNet-1K with improved accuracy and reduced time to train
- PatrickStar: Parallel Training of Pre-trained Models via Chunk-based Memory Management
- OneFlow: Redesign the Distributed Deep Learning Framework from Scratch
- Towards Better Accuracy-efficiency Trade-offs: Divide and Co-training
- Hanayo: Harnessing Wave-like Pipeline Parallelism for Enhanced Large Model Training Efficiency
- Tesseract: Parallelize the Tensor Parallelism Efficiently
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine Translation
- QuartzNet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions
- UNIT: Unifying Tensorized Instruction Compilation
- SPRING: A Sparsity-Aware Reduced-Precision Monolithic 3D CNN Accelerator Architecture for Training and Inference
- CNN Filter DB: An Empirical Investigation of Trained Convolutional Filters
- Rethinking "Batch" in BatchNorm
- Training DNNs with Hybrid Block Floating Point
- Efficient Quantized Sparse Matrix Operations on Tensor Cores
- Deep Volumetric Ambient Occlusion
- DS6, Deformation-aware Semi-supervised Learning: Application to Small Vessel Segmentation with Noisy Training Data
- Data Movement Is All You Need: A Case Study on Optimizing Transformers
- Towards Understanding Fast Adversarial Training
- On Extractive and Abstractive Neural Document Summarization with Transformer Language Models
- Adaptive Regularization of Labels
- EnforceSNN: Enabling Resilient and Energy-Efficient Spiking Neural Network Inference considering Approximate DRAMs for Embedded Systems
- TF-Replicator: Distributed Machine Learning for Researchers
- QGAN: Quantized Generative Adversarial Networks
- JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus
- Boosting EfficientNets Ensemble Performance via Pseudo-Labels and Synthetic Images by pix2pixHD for Infection and Ischaemia Classification in Diabetic Foot Ulcers
- Degree-Quant: Quantization-Aware Training for Graph Neural Networks
- Revisiting Neural Retrieval on Accelerators
- On the Utility of Gradient Compression in Distributed Training Systems
- Differentiable Model Compression via Pseudo Quantization Noise
- 1st Place Solution for Waymo Open Dataset Challenge -- 3D Detection and Domain Adaptation
- Local Critic Training for Model-Parallel Learning of Deep Neural Networks
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- IGUANe: a 3D generalizable CycleGAN for multicenter harmonization of brain MR images
- Deep Learning as a Mixed Convex-Combinatorial Optimization Problem
- Mixed-precision explicit stabilized Runge-Kutta methods for single- and multi-scale differential equations
- Backprop with Approximate Activations for Memory-efficient Network Training
- Blockwise Self-Attention for Long Document Understanding
- Time-aware Large Kernel Convolutions
- Post-Training Piecewise Linear Quantization for Deep Neural Networks
- Post-Training 4-bit Quantization on Embedding Tables
- MatchboxNet: 1D Time-Channel Separable Convolutional Neural Network Architecture for Speech Commands Recognition
- Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
- The Power of Selecting Key Blocks with Local Pre-ranking for Long Document Information Retrieval
- Pre-Trained Models: Past, Present and Future
- ZEN: Pre-training Chinese Text Encoder Enhanced by N-gram Representations
- DaCapo: Accelerating Continuous Learning in Autonomous Systems for Video Analytics
- Dual-path Self-Attention RNN for Real-Time Speech Enhancement
- Finding Nano-Ötzi: Semi-Supervised Volume Visualization for Cryo-Electron Tomography
- Efficient 3D Fully Convolutional Networks for Pulmonary Lobe Segmentation in CT Images
- HPC AI500: The Methodology, Tools, Roofline Performance Models, and Metrics for Benchmarking HPC AI Systems
- LIT: Block-wise Intermediate Representation Training for Model Compression
- TalkNet: Fully-Convolutional Non-Autoregressive Speech Synthesis Model
- RecJPQ: Training Large-Catalogue Sequential Recommenders
- Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
- FedDCT: Federated Learning of Large Convolutional Neural Networks on Resource Constrained Devices using Divide and Collaborative Training
- Fix your classifier: the marginal value of training the last weight layer
- FxP-QNet: A Post-Training Quantizer for the Design of Mixed Low-Precision DNNs with Dynamic Fixed-Point Representation
- CPT: Efficient Deep Neural Network Training via Cyclic Precision
- PDE-Driven Spatiotemporal Disentanglement
- Low-Precision Reinforcement Learning: Running Soft Actor-Critic in Half Precision
- MXR-U-Nets for Real Time Hyperspectral Reconstruction
- An Efficient Transformer Decoder with Compressed Sub-layers
- PDPU: An Open-Source Posit Dot-Product Unit for Deep Learning Applications
- Towards recognizing the light facet of the Higgs Boson
- Improving Neural Network Training using Dynamic Learning Rate Schedule for PINNs and Image Classification
- KAISA: An Adaptive Second-Order Optimizer Framework for Deep Neural Networks
- Training High-Performance and Large-Scale Deep Neural Networks with Full 8-bit Integers
- COLD: Towards the Next Generation of Pre-Ranking System
- A Fast and Robust BERT-based Dialogue State Tracker for Schema-Guided Dialogue Dataset
- ActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed Training
- Memory Optimization for Deep Networks
- Training Multilingual Pre-trained Language Model with Byte-level Subwords
- Meta Batch-Instance Normalization for Generalizable Person Re-Identification
- ScaDLES: Scalable Deep Learning over Streaming data at the Edge
- Large-Scale Deep Learning Optimizations: A Comprehensive Survey
- Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
- ESPnet-ST: All-in-One Speech Translation Toolkit
- LazyFormer: Self Attention with Lazy Update
- Deep Learning Training on the Edge with Low-Precision Posits
- Offline Handwritten Chinese Text Recognition with Convolutional Neural Networks
- Real-time Person Re-identification at the Edge: A Mixed Precision Approach
- EventGraD: Event-Triggered Communication in Parallel Machine Learning
- Organ Segmentation From Full-size CT Images Using Memory-Efficient FCN
- Accelerating CNN Training by Pruning Activation Gradients
- Revisiting BFloat16 Training
- Faster Neural Network Training with Approximate Tensor Operations
- Sharing Attention Weights for Fast Transformer
- Multi-node Bert-pretraining: Cost-efficient Approach
- word2ket: Space-efficient Word Embeddings inspired by Quantum Entanglement
- Rethinking Floating Point Overheads for Mixed Precision DNN Accelerators
- Stable, Fast and Accurate: Kernelized Attention with Relative Positional Encoding
- Deep Neural Networks to Correct Sub-Precision Errors in CFD
- AuGPT: Auxiliary Tasks and Data Augmentation for End-To-End Dialogue with Pre-Trained Language Models
- NeurST: Neural Speech Translation Toolkit
- AI Enabling Technologies: A Survey
- WrapNet: Neural Net Inference with Ultra-Low-Resolution Arithmetic
- LAMP: Large Deep Nets with Automated Model Parallelism for Image Segmentation
- LazyTensor: combining eager execution with domain-specific compilers
- Revitalizing CNN Attentions via Transformers in Self-Supervised Visual Representation Learning
- LICHEE: Improving Language Model Pre-training with Multi-grained Tokenization
- Rethinking Training from Scratch for Object Detection
- Boundary-Refined Prototype Generation: A General End-to-End Paradigm for Semi-Supervised Semantic Segmentation
- Accelerating Distributed ML Training via Selective Synchronization
- GraVAC: Adaptive Compression for Communication-Efficient Distributed DL Training
- Simplified TinyBERT: Knowledge Distillation for Document Retrieval
- Training Deep Neural Networks Using Posit Number System
- Benchmarking the Performance and Energy Efficiency of AI Accelerators for AI Training
- Demystifying the MLPerf Benchmark Suite
- Characterizing Deep Learning Training Workloads on Alibaba-PAI
- HG-Caffe: Mobile and Embedded Neural Network GPU (OpenCL) Inference Engine with FP16 Supporting
- GAN Cocktail: mixing GANs without dataset access
- Learning compositional functions via multiplicative weight updates
- Neural Network Libraries: A Deep Learning Framework Designed from Engineers' Perspectives
- TalkNet 2: Non-Autoregressive Depth-Wise Separable Convolutional Model for Speech Synthesis with Explicit Pitch and Duration Prediction
- Dynamic Runtime Feature Map Pruning
- SimTriplet: Simple Triplet Representation Learning with a Single GPU
- Enhanced 3D Brain Tumor Segmentation Using Assorted Precision Training
- Towards Rapid and Robust Adversarial Training with One-Step Attacks
- FlexSA: Flexible Systolic Array Architecture for Efficient Pruned DNN Model Training
- Conditional Image Generation and Manipulation for User-Specified Content
- PyTorchDIA: A flexible, GPU-accelerated numerical approach to Difference Image Analysis
- QuTiBench: Benchmarking Neural Networks on Heterogeneous Hardware
- 8-bit Optimizers via Block-wise Quantization
- fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit
- Self-Supervised GAN Compression
- Climate Modelling in Low-Precision: Effects of both Deterministic & Stochastic Rounding
- GLSD: The Global Large-Scale Ship Database and Baseline Evaluations
- A Simple but Effective BERT Model for Dialog State Tracking on Resource-Limited Systems
- Mesa: A Memory-saving Training Framework for Transformers
- Training Deep Neural Network in Limited Precision
- Leveraging Mixed Precision in Exponential Time Integration Methods
- Accelerating Gravitational -Body Simulations Using the RISC-V-Based Tenstorrent Wormhole
- Compute, Time and Energy Characterization of Encoder-Decoder Networks with Automatic Mixed Precision Training
- Taking Notes on the Fly Helps BERT Pre-training
- SGD-QA: Fast Schema-Guided Dialogue State Tracking for Unseen Services
- EfficientQA : a RoBERTa Based Phrase-Indexed Question-Answering System
- A Runtime-Based Computational Performance Predictor for Deep Neural Network Training
- Split to Be Slim: An Overlooked Redundancy in Vanilla Convolution
- Analyzing Machine Learning Workloads Using a Detailed GPU Simulator
- Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training
- Hydra: A Peer to Peer Distributed Training & Data Collection Framework
- Revisiting Language Encoding in Learning Multilingual Representations
- Layered gradient accumulation and modular pipeline parallelism: fast and efficient training of large language models
- Pagsusuri ng RNN-based Transfer Learning Technique sa Low-Resource Language
- Distributed Low Precision Training Without Mixed Precision
- Contextual embedding and model weighting by fusing domain knowledge on Biomedical Question Answering
- An FPGA Accelerated Method for Training Feed-forward Neural Networks Using Alternating Direction Method of Multipliers and LSMR
- Generalisation of Cyberbullying Detection
- COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy Compression
- Multilingual Transformers for Product Matching -- Experiments and a New Benchmark in Polish
- A Multi-site Study of a Breast Density Deep Learning Model for Full-field Digital Mammography Images and Synthetic Mammography Images
- ZeroGrad : Mitigating and Explaining Catastrophic Overfitting in FGSM Adversarial Training
- Triple-cooperative Video Shadow Detection
- Scaling Law for Recommendation Models: Towards General-purpose User Representations
- LEWIS: Levenshtein Editing for Unsupervised Text Style Transfer
- Attention-based fusion of semantic boundary and non-boundary information to improve semantic segmentation
- Performance Analysis of Deep Learning Workloads on Leading-edge Systems
- The Complexity Dynamics of Grokking
- A Simple Non-i.i.d. Sampling Approach for Efficient Training and Better Generalization
- How Low Can We Go: Trading Memory for Error in Low-Precision Training
- Accelerated CNN Training Through Gradient Approximation
- A Protection Method of Trained CNN Model Using Feature Maps Transformed With Secret Key From Unauthorized Access
- NTIRE 2020 Challenge on Spectral Reconstruction from an RGB Image
- Enabling Binary Neural Network Training on the Edge
- Dynamic Efficient Adversarial Training Guided by Gradient Magnitude
- Yield Loss Reduction and Test of AI and Deep Learning Accelerators
- Partition and Code: learning how to compress graphs
- In-training Matrix Factorization for Parameter-frugal Neural Machine Translation
- Hessian-aware Quantized Node Embeddings for Recommendation
- Learning In Practice: Reasoning About Quantization
- On the Downstream Performance of Compressed Word Embeddings
- Solving Traffic4Cast Competition with U-Net and Temporal Domain Adaptation
- Robust Single-step Adversarial Training with Regularizer
- OD-SGD: One-step Delay Stochastic Gradient Descent for Distributed Training
- Mini-batch Serialization: CNN Training with Inter-layer Data Reuse
- Overfitting or Underfitting? Understand Robustness Drop in Adversarial Training
- Faster object tracking pipeline for real time tracking
- Mixed precision in Graphics Processing Unit
- NTT's Machine Translation Systems for WMT19 Robustness Task
- Synthesizing Compact Hardware for Accelerating Inference from Physical Signals in Sensors
- Attentive fine-tuning of Transformers for Translation of low-resourced languages @LoResMT 2021
- Scaling Analog Photonic Accelerators for Byte-Size, Integer General Matrix Multiply (GEMM) Kernels
- Floating-Point Neural Networks Are Provably Robust Universal Approximators
- Progressive Compressed Records: Taking a Byte out of Deep Learning Data
- MIX : a Multi-task Learning Approach to Solve Open-Domain Question Answering
- DDU-Net: A Domain Decomposition-Based CNN for High-Resolution Image Segmentation on Multiple GPUs
- MG-WFBP: Merging Gradients Wisely for Efficient Communication in Distributed Deep Learning
- Optimizing High-Dimensional Physics Simulations via Composite Bayesian Optimization
- Penetrating the Fog: the Path to Efficient CNN Models
- Accelerating Deep Learning with Dynamic Data Pruning
- Transfer Learning with Binary Neural Networks
- Training for Speech Recognition on Coprocessors
- MoRS: An Approximate Fault Modelling Framework for Reduced-Voltage SRAMs
- Towards Fully 8-bit Integer Inference for the Transformer Model
- LightSeq2: Accelerated Training for Transformer-based Models on GPUs
- Initializing Perturbations in Multiple Directions for Fast Adversarial Training
- OpTorch: Optimized deep learning architectures for resource limited environments
- Low-Complexity LSTM Training and Inference with FloatSD8 Weight Representation
- Weight Distillation: Transferring the Knowledge in Neural Network Parameters
- Deep Neural Network Training without Multiplications
- k-Same-Siamese-GAN: k-Same Algorithm with Generative Adversarial Network for Facial Image De-identification with Hyperparameter Tuning and Mixed Precision Training
- Empirical Evaluation of Deep Learning Model Compression Techniques on the WaveNet Vocoder