The Computational Limits of Deep Learning
arXiv:2007.05558
Abstract
Deep learning's recent history has been one of achievement: from triumphing over humans in the game of Go to world-leading performance in image classification, voice recognition, translation, and other tasks. But this progress has come with a voracious appetite for computing power. This article catalogs the extent of this dependency, showing that progress across a wide variety of applications is strongly reliant on increases in computing power. Extrapolating forward this reliance reveals that progress along current lines is rapidly becoming economically, technically, and environmentally unsustainable. Thus, continued progress in these applications will require dramatically more computationally-efficient methods, which will either have to come from changes to deep learning or from moving to other machine learning methods.
33 pages, 8 figures
Cited by in corpus (18)
- Compute Trends Across Three Eras of Machine Learning
- Traffic Prediction using Artificial Intelligence: Review of Recent Advances and Emerging Opportunities
- Compute and Energy Consumption Trends in Deep Learning Inference
- Bottom-up and top-down approaches for the design of neuromorphic processing systems: Tradeoffs and synergies between natural and artificial intelligence
- Great Power, Great Responsibility: Recommendations for Reducing Energy for Training Language Models
- Underwater Acoustic Communication Channel Modeling using Reservoir Computing
- PositNN: Training Deep Neural Networks with Mixed Low-Precision Posit
- ETLP: Event-based Three-factor Local Plasticity for online learning with neuromorphic hardware
- MUCM-Net: A Mamba Powered UCM-Net for Skin Lesion Segmentation
- Online Deterministic Annealing for Classification and Clustering
- Annealing Optimization for Progressive Learning with Stochastic Approximation
- Sub-milliwatt threshold power and tunable-bias all-optical nonlinear activation function using vanadium dioxide for wavelength-division multiplexing photonic neural networks
- Augmenting semantic lexicons using word embeddings and transfer learning
- tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
- No Free Lunch: Balancing Learning and Exploitation at the Network Edge
- ExplainFix: Explainable Spatially Fixed Deep Networks
- Thrill-K Architecture: Towards a Solution to the Problem of Knowledge Based Understanding
- The Tensor Track VII: From Quantum Gravity to Artificial Intelligence