Full error analysis for the training of deep neural networks
arXiv:1910.00121 · doi:10.1142/S021902572150020X
Abstract
Deep learning algorithms have been applied very successfully in recent years to a range of problems out of reach for classical solution paradigms. Nevertheless, there is no completely rigorous mathematical error and convergence analysis which explains the success of deep learning algorithms. The error of a deep learning algorithm can in many situations be decomposed into three parts, the approximation error, the generalization error, and the optimization error. In this work we estimate for a certain deep learning algorithm each of these three errors and combine these three error estimates to obtain an overall error analysis for the deep learning algorithm under consideration. In particular, we thereby establish convergence with a suitable convergence speed for the overall error of the deep learning algorithm under consideration. Our convergence speed analysis is far from optimal and the convergence speed that we establish is rather slow, increases exponentially in the dimensions, and, in particular, suffers from the curse of dimensionality. The main contribution of this work is, instead, to provide a full error analysis (i) which covers each of the three different sources of errors usually emerging in deep learning algorithms and (ii) which merges these three sources of errors into one overall error estimate for the considered deep learning algorithm.
53 pages
References in corpus (3)
Cited by in corpus (14)
- A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions
- On the existence of global minima and convergence analyses for gradient descent methods in the training of deep neural networks
- A proof of convergence for stochastic gradient descent in the training of artificial neural networks with ReLU activation for constant target functions
- The gap between theory and practice in function approximation with deep neural networks
- Learning the random variables in Monte Carlo simulations with stochastic gradient descent: Machine learning for parametric PDEs and financial derivative pricing
- Partition of unity networks: deep hp-approximation
- Deep Neural Networks Are Effective At Learning High-Dimensional Hilbert-Valued Functions From Limited Data
- High-dimensional approximation spaces of artificial neural networks and applications to partial differential equations
- A Neural Solver for Variational Problems on CAD Geometries with Application to Electric Machine Simulation
- Convergence rates for gradient descent in the training of overparameterized artificial neural networks with piecewise affine activation
- On the approximation of functions by tanh neural networks
- Strong overall error analysis for the training of artificial neural networks via random initializations
- Polynomial-Spline Neural Networks with Exact Integrals
- Multilevel Monte Carlo learning