Low-Memory Neural Network Training: A Technical Report
arXiv:1904.10631
Abstract
Memory is increasingly often the bottleneck when training neural network models. Despite this, techniques to lower the overall memory requirements of training have been less widely studied compared to the extensive literature on reducing the memory requirements of inference. In this paper we study a fundamental question: How much memory is actually needed to train a neural network? To answer this question, we profile the overall memory usage of training on two representative deep learning benchmarks -- the WideResNet model for image classification and the DynamicConv Transformer model for machine translation -- and comprehensively evaluate four standard techniques for reducing the training memory requirements: (1) imposing sparsity on the model, (2) using low precision, (3) microbatching, and (4) gradient checkpointing. We explore how each of these techniques in isolation affects both the peak memory usage of training and the quality of the end model, and explore the memory, accuracy, and computation tradeoffs incurred when combining these techniques. Using appropriate combinations of these techniques, we show that it is possible to the reduce the memory required to train a WideResNet-28-2 on CIFAR-10 by up to 60.7x with a 0.4% loss in accuracy, and reduce the memory required to train a DynamicConv model on IWSLT'14 German to English translation by up to 8.7x with a BLEU score drop of 0.15.
Version notes: Copyedits and citation fixes
References in corpus (28)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Mixed Precision Training
- Trained Ternary Quantization
- To prune, or not to prune: exploring the efficacy of pruning for model compression
- Compressing Neural Networks with the Hashing Trick
- The State of Sparsity in Deep Neural Networks
- Revisiting Small Batch Training for Deep Neural Networks
- Pay Less Attention with Lightweight and Dynamic Convolutions
- Batch Renormalization: Towards Reducing Minibatch Dependence in Batch-Normalized Models
- Scalable Methods for 8-bit Training of Neural Networks
- The Reversible Residual Network: Backpropagation Without Storing Activations
- Deep Rewiring: Training very sparse deep networks
- Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
- SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks
- Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization
- Fixup Initialization: Residual Learning Without Normalization
- Trained Quantization Thresholds for Accurate and Efficient Fixed-Point Inference of Deep Neural Networks
- Optimizing and Visualizing Deep Learning for Benign/Malignant Classification in Breast Tumors
- HAQ: Hardware-Aware Automated Quantization with Mixed Precision
- Learning Fast Algorithms for Linear Transforms Using Butterfly Factorizations
- Traditional and Heavy-Tailed Self Regularization in Neural Network Models
- Streaming Normalization: Towards Simpler and More Biologically-plausible Normalizations for Online and Recurrent Learning
- Escoin: Efficient Sparse Convolutional Neural Network Inference on GPUs
- Memory-Efficient Adaptive Optimization
- Sparse evolutionary Deep Learning with over one million artificial neurons on commodity hardware
Cited by in corpus (16)
- Reformer: The Efficient Transformer
- Drawing Early-Bird Tickets: Towards More Efficient Training of Deep Networks
- Zero-shot Entity Linking with Efficient Long Range Sequence Modeling
- Blockwise Self-Attention for Long Document Understanding
- On using distributed representations of source code for the detection of C security vulnerabilities
- Campfire: Compressible, Regularization-Free, Structured Sparse Training for Hardware Accelerators
- Minimizing FLOPs to Learn Efficient Sparse Representations
- word2ket: Space-efficient Word Embeddings inspired by Quantum Entanglement
- On improving deep learning generalization with adaptive sparse connectivity
- How Low Can We Go: Trading Memory for Error in Low-Precision Training
- Dynamic Tensor Rematerialization
- On the Downstream Performance of Compressed Word Embeddings
- Improving Formality Style Transfer with Context-Aware Rule Injection
- Enabling Binary Neural Network Training on the Edge
- Hydra: A System for Large Multi-Model Deep Learning
- Enabling Incremental Training with Forward Pass for Edge Devices