Compute and Energy Consumption Trends in Deep Learning Inference
arXiv:2109.05472 · doi:10.1016/j.suscom.2023.100857
Abstract
The progress of some AI paradigms such as deep learning is said to be linked to an exponential growth in the number of parameters. There are many studies corroborating these trends, but does this translate into an exponential increase in energy consumption? In order to answer this question we focus on inference costs rather than training costs, as the former account for most of the computing effort, solely because of the multiplicative factors. Also, apart from algorithmic innovations, we account for more specific and powerful hardware (leading to higher FLOPS) that is usually accompanied with important energy efficiency optimisations. We also move the focus from the first implementation of a breakthrough paper towards the consolidated version of the techniques one or two year later. Under this distinctive and comprehensive perspective, we study relevant models in the areas of computer vision and natural language processing: for a sustained increase in performance we see a much softer growth in energy consumption than previously anticipated. The only caveat is, yet again, the multiplicative factor, as future AI increases penetration and becomes more pervasive.
For a revised version and its published version refer to: Desislavov, Radosvet, Fernando Martínez-Plumed, and José Hernández-Orallo. Trends in AI inference energy consumption: Beyond the performance-vs-parameter laws of deep learning. Sustainable Computing: Informatics and Systems, Volume 38, April 2023. (https://doi.org/10.1016/j.suscom.2023.100857)
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- On the Opportunities and Risks of Foundation Models
- Scaling Laws for Neural Language Models
- EfficientNetV2: Smaller Models and Faster Training
- Deep Learning Scaling is Predictable, Empirically
- Billion-scale semi-supervised learning for image classification
- The Computational Limits of Deep Learning
- High-Performance Large-Scale Image Recognition Without Normalization
- Scaling Laws for Autoregressive Generative Modeling
- Measuring the Algorithmic Efficiency of Neural Networks
- Great Power, Great Responsibility: Recommendations for Reducing Energy for Training Language Models
Cited by in corpus (20)
- All-optical image denoising using a diffractive visual processor
- Unveiling Energy Efficiency in Deep Learning: Measurement, Prediction, and Scoring across Edge Devices
- A Survey on AI-driven Energy Optimisation in Terrestrial Next Generation Radio Access Networks
- Energy and Carbon Considerations of Fine-Tuning BERT
- Efficiency is Not Enough: A Critical Perspective of Environmentally Sustainable AI
- Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
- Sustainable Edge Intelligence Through Energy-Aware Early Exiting
- FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems
- BLAZE: Cross-Language and Cross-Project Bug Localization via Dynamic Chunking and Hard Example Learning
- Identifying architectural design decisions for achieving green ML serving
- Missing Data as Augmentation in the Earth Observation Domain: A Multi-View Learning Approach
- Accelerating the drive towards energy-efficient generative AI with quantum computing algorithms
- Hyperdimensional Representation Learning for Node Classification and Link Prediction
- Insights into resource utilization of code small language models serving with runtime engines and execution providers
- acoupi: An Open-Source Python Framework for Deploying Bioacoustic AI Models on Edge Devices
- Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
- Energy Efficiency in AI for 5G and Beyond: A DeepRx Case Study
- Empirically-Calibrated H100 Node Power Models for Reducing Uncertainty in AI Training Energy Estimation
- MRM3: Machine Readable ML Model Metadata
- Phoeni6: a Systematic Approach for Evaluating the Energy Consumption of Neural Networks