Data Movement Is All You Need: A Case Study on Optimizing Transformers
arXiv:2007.00072
Abstract
Transformers are one of the most important machine learning workloads today. Training one is a very compute-intensive task, often taking days or weeks, and significant attention has been given to optimizing transformers. Despite this, existing implementations do not efficiently utilize GPUs. We find that data movement is the key bottleneck when training. Due to Amdahl's Law and massive improvements in compute performance, training has now become memory-bound. Further, existing frameworks use suboptimal data layouts. Using these insights, we present a recipe for globally optimizing data movement in transformers. We reduce data movement by up to 22.91% and overall achieve a 1.30x performance improvement over state-of-the-art frameworks when training a BERT encoder layer and 1.19x for the entire BERT. Our approach is applicable more broadly to optimizing deep neural networks, and offers insight into how to tackle emerging performance bottlenecks.
22 pages, 8 figures; MLSys 2021 camera ready
References in corpus (18)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Language Models are Few-Shot Learners
- Scaling Laws for Neural Language Models
- Sequence Transduction with Recurrent Neural Networks
- cuDNN: Efficient Primitives for Deep Learning
- Generating Long Sequences with Sparse Transformers
- Reformer: The Efficient Transformer
- Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes
- Device Placement Optimization with Reinforcement Learning
- MLIR: A Compiler Infrastructure for the End of Moore's Law
- Intel nGraph: An Intermediate Representation, Compiler, and Executor for Deep Learning
- Dissecting the Graphcore IPU Architecture via Microbenchmarking
- Sparse Sinkhorn Attention
- Transformer-Transducer: End-to-End Speech Recognition with Self-Attention
- Fast Transformer Decoding: One Write-Head is All You Need
- TF-Replicator: Distributed Machine Learning for Researchers
- TIRAMISU: A Polyhedral Compiler for Dense and Sparse Deep Learning
Cited by in corpus (6)
- Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
- SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on Mobile
- Demystifying BERT: Implications for Accelerator Design
- MCUBERT: Memory-Efficient BERT Inference on Commodity Microcontrollers
- FastSeq: Make Sequence Generation Faster
- Pebbles, Graphs, and a Pinch of Combinatorics: Towards Tight I/O Lower Bounds for Statically Analyzable Programs