The Unreasonable Effectiveness of Deep Learning in Artificial Intelligence
arXiv:2002.04806 · doi:10.1073/pnas.1907373117
Abstract
Deep learning networks have been trained to recognize speech, caption photographs and translate text between languages at high levels of performance. Although applications of deep learning networks to real world problems have become ubiquitous, our understanding of why they are so effective is lacking. These empirical results should not be possible according to sample complexity in statistics and non-convex optimization theory. However, paradoxes in the training and effectiveness of deep learning networks are being investigated and insights are being found in the geometry of high-dimensional spaces. A mathematical theory of deep learning would illuminate how they function, allow us to assess the strengths and weaknesses of different network architectures and lead to major improvements. Deep learning has provided natural ways for humans to communicate with digital devices and is foundational for building artificial general intelligence. Deep learning was inspired by the architecture of the cerebral cortex and insights into autonomy and general intelligence may be found in other brain regions that are essential for planning and survival, but major breakthroughs will be needed to achieve these goals.
References in corpus (2)
Cited by in corpus (20)
- Machine Learning and Deep Learning -- A review for Ecologists
- Large Language Models and the Reverse Turing Test
- Optimized ensemble deep learning framework for scalable forecasting of dynamics containing extreme events
- A comprehensive and fair comparison of two neural operators (with practical extensions) based on FAIR data
- SoftHebb: Bayesian Inference in Unsupervised Hebbian Soft Winner-Take-All Networks
- Deep Learning for Scene Classification: A Survey
- Open Problems in Cooperative AI
- Deep learning for nano-photonic materials -- The solution to everything!?
- Single Unit Status in Deep Convolutional Neural Network Codes for Face Identification: Sparseness Redefined
- Should attention be all we need? The epistemic and ethical implications of unification in machine learning
- Machine Learning for Robust Identification of Complex Nonlinear Dynamical Systems: Applications to Earth Systems Modeling
- Deep Neural Networks for Automatic Speaker Recognition Do Not Learn Supra-Segmental Temporal Features
- Polynomial Ridge Flowfield Estimation
- Unified Regularity Measures for Sample-wise Learning and Generalization
- Modern Machine and Deep Learning Systems as a way to achieve Man-Computer Symbiosis
- Embed Everything: A Method for Efficiently Co-Embedding Multi-Modal Spaces
- NUTS, NARS, and Speech
- Bayesian Deep Learning Hyperparameter Search for Robust Function Mapping to Polynomials with Noise
- Comments on Sejnowski's "The unreasonable effectiveness of deep learning in artificial intelligence" [arXiv:2002.04806]
- A brain basis of dynamical intelligence for AI and computational neuroscience