The Modern Mathematics of Deep Learning
arXiv:2105.04026 · doi:10.1017/9781009025096.002
Abstract
We describe the new field of mathematical analysis of deep learning. This field emerged around a list of research questions that were not answered within the classical framework of learning theory. These questions concern: the outstanding generalization power of overparametrized neural networks, the role of depth in deep architectures, the apparent absence of the curse of dimensionality, the surprisingly successful optimization performance despite the non-convexity of the problem, understanding what features are learned, why deep architectures perform exceptionally well in physical problems, and which fine aspects of an architecture affect the behavior of a learning task in which way. We present an overview of modern approaches that yield partial answers to these questions. For selected approaches, we describe the main ideas in more detail.
A version of this review paper appears as a chapter in the book "Mathematical Aspects of Deep Learning" by Cambridge University Press
References in corpus (29)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Deep Learning in Neural Networks: An Overview
- Neural Architecture Search with Reinforcement Learning
- The Loss Surfaces of Multilayer Networks
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Exploring Generalization in Deep Learning
- Inhomogeneous backflow transformations in quantum Monte Carlo calculations
- Escaping From Saddle Points --- Online Stochastic Gradient for Tensor Decomposition
- Spectrally-normalized margin bounds for neural networks
- Interactions between Large Molecules: Puzzle for Reference Quantum-Mechanical Methods
- Fantastic Generalization Measures and Where to Find Them
- Why Deep Neural Networks for Function Approximation?
- In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning
- Working Locally Thinking Globally: Theoretical Guarantees for Convolutional Sparse Coding
- Norm-Based Capacity Control in Neural Networks
- Towards a Mathematical Understanding of Neural Network-Based Machine Learning: what we know and what we don't
- Theory of Deep Learning III: explaining the non-overfitting puzzle
- Multiple Descent: Design Your Own Generalization Curve
- Exponential ReLU Neural Network Approximation Rates for Point and Edge Singularities
- What causes the test error? Going beyond bias-variance via ANOVA
- Expressivity of Deep Neural Networks
- On the Banach spaces associated with multi-layer ReLU networks: Function representation, approximation theory and gradient descent dynamics
- Sparse Neural Networks Topologies
- Neural network approximation and estimation of classifiers with classification boundary in a Barron class
- Proof of the Theory-to-Practice Gap in Deep Learning via Sampling Complexity bounds for Neural Network Approximation Spaces
- A priori estimates for classification problems using neural networks
- Solving the electronic Schrödinger equation for multiple nuclear geometries with weight-sharing deep neural networks
- Elementary superexpressive activations
- Deep neural network approximation for high-dimensional parabolic Hamilton-Jacobi-Bellman equations
Cited by in corpus (17)
- Incorporating Domain Knowledge into Deep Neural Networks
- A Review of Some Techniques for Inclusion of Domain-Knowledge into Deep Neural Networks
- Combining physics-based and data-driven models: advancing the frontiers of research with Scientific Machine Learning
- Error convergence and engineering-guided hyperparameter search of PINNs: towards optimized I-FENN performance
- Limitations of Deep Learning for Inverse Problems on Digital Hardware
- Learning-informed parameter identification in nonlinear time-dependent PDEs
- Emergence of Concepts in DNNs?
- Towards a population-informed approach to the definition of data-driven models for structural dynamics
- Random feature neural networks learn Black-Scholes type PDEs without curse of dimensionality
- On generalization bounds for deep networks based on loss surface implicit regularization
- Rosenblatt's first theorem and frugality of deep learning
- A machine learning approach to predict L-edge x-ray absorption spectra of light transition metal ion compounds
- Neural variance reduction for stochastic differential equations
- A Low Rank Neural Representation of Entropy Solutions
- ReLU Neural Networks of Polynomial Size for Exact Maximum Flow Computation
- Neural Networks in Fréchet spaces
- First passage time in space-dependent stochastic resetting