Compute Trends Across Three Eras of Machine Learning
arXiv:2202.05924 · doi:10.1109/IJCNN55064.2022.9891914
Abstract
Compute, data, and algorithmic advances are the three fundamental factors that guide the progress of modern Machine Learning (ML). In this paper we study trends in the most readily quantified factor - compute. We show that before 2010 training compute grew in line with Moore's law, doubling roughly every 20 months. Since the advent of Deep Learning in the early 2010s, the scaling of training compute has accelerated, doubling approximately every 6 months. In late 2015, a new trend emerged as firms developed large-scale ML models with 10 to 100-fold larger requirements in training compute. Based on these observations we split the history of compute in ML into three eras: the Pre Deep Learning Era, the Deep Learning Era and the Large-Scale Era. Overall, our work highlights the fast-growing compute requirements for training advanced ML systems.
References in corpus (20)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Improving neural networks by preventing co-adaptation of feature detectors
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Generative Adversarial Networks
- Neural Architecture Search with Reinforcement Learning
- PaLM: Scaling Language Modeling with Pathways
- Scaling Laws for Neural Language Models
- Flamingo: a Visual Language Model for Few-Shot Learning
- DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker
- LaMDA: Language Models for Dialog Applications
- Training Compute-Optimal Large Language Models
- Deep Learning Scaling is Predictable, Empirically
- Deep Learning Recommendation Model for Personalization and Recommendation Systems
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Self-supervised Pretraining of Visual Features in the Wild
- The Cost of Training NLP Models: A Concise Overview
- PanGu-: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation
- Yuan 1.0: Large-Scale Pre-trained Language Model in Zero-Shot and Few-Shot Learning
- Scaling Scaling Laws with Board Games
Cited by in corpus (30)
- Experimentally realized in situ backpropagation for deep learning in nanophotonic neural networks
- How Generative AI models such as ChatGPT can be (Mis)Used in SPC Practice, Education, and Research? An Exploratory Study
- Large Language Models and the Reverse Turing Test
- ChatGPT Needs SPADE (Sustainability, PrivAcy, Digital divide, and Ethics) Evaluation: A Review
- Quantum Machine Learning: from physics to software engineering
- Astronomia ex machina: a history, primer, and outlook on neural networks in astronomy
- Carbon Footprint of Selecting and Training Deep Learning Models for Medical Image Analysis
- Three lines of defense against risks from AI
- Trends in Energy Estimates for Computing in AI/Machine Learning Accelerators, Supercomputers, and Compute-Intensive Applications
- Will Code Remain a Relevant User Interface for End-User Programming with Generative AI Models?
- Efficiency is Not Enough: A Critical Perspective of Environmentally Sustainable AI
- EC-NAS: Energy Consumption Aware Tabular Benchmarks for Neural Architecture Search
- Ultrafast Coherent Dynamics of Microring Modulators
- Rethinking model prototyping through the MedMNIST+ dataset collection
- Operating critical machine learning models in resource constrained regimes
- A Green(er) World for A.I
- Phase-space analysis of a two-section InP laser as an all-optical spiking neuron: dependency on control and design parameters
- Energy Efficiency trends in HPC: what high-energy and astrophysicists need to know
- A 262 TOPS Hyperdimensional Photonic AI Accelerator powered by a Si3N4 microcomb laser
- FedMoE-DA: Federated Mixture of Experts via Domain Aware Fine-grained Aggregation
- Activation Compression of Graph Neural Networks using Block-wise Quantization with Improved Variance Minimization
- A Reconfigurable Stream-Based FPGA Accelerator for Bayesian Confidence Propagation Neural Networks
- Experimental Standards for Deep Learning in Natural Language Processing Research
- TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
- Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
- How Small is Big Enough? Open Labeled Datasets and the Development of Deep Learning
- Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
- A Memory Efficient Adjoint Method to Enable Billion Parameter Optimization on a Single GPU in Dynamic Problems
- Accelerating Sharded Data Parallelism at Scale with Federated Learning
- Oxide Interface-Based Polymorphic Electronic Devices for Neuromorphic Computing