26 citations · 40 across the 4 of their papers we have counts for
6 papers
Efficient Sobolev approximation of linear parabolic PDEs in high dimensions
Patrick Cheridito, Florian Rossmannek
In this paper, we study the error in first order Sobolev norm in the approximation of solutions to linear parabolic PDEs. We use a Monte Carlo Euler scheme obtained from combining…
Gradient descent provably escapes saddle points in the training of shallow ReLU networks
Patrick Cheridito, Arnulf Jentzen, Florian Rossmannek
Dynamical systems theory has recently been applied in optimization to prove that gradient descent algorithms bypass so-called strict saddle points of the loss function. However, in…
Landscape analysis for shallow neural networks: complete classification of critical points for affine target functions
Patrick Cheridito, Arnulf Jentzen, Florian Rossmannek
In this paper, we analyze the landscape of the true loss of neural networks with one hidden layer and ReLU, leaky ReLU, or quadratic activation. In all three cases, we provide a co…
A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions
Patrick Cheridito, Arnulf Jentzen, Adrian Riekert +1
Gradient descent optimization algorithms are the standard ingredients that are used to train artificial neural networks (ANNs). Even though a huge number of numerical simulations i…
Non-convergence of stochastic gradient descent in the training of deep neural networks
Patrick Cheridito, Arnulf Jentzen, Florian Rossmannek
Deep neural networks have successfully been trained in various application areas with stochastic gradient descent. However, there exists no rigorous mathematical explanation why th…
Efficient approximation of high-dimensional functions with neural networks
Patrick Cheridito, Arnulf Jentzen, Florian Rossmannek
In this paper, we develop a framework for showing that neural networks can overcome the curse of dimensionality in different high-dimensional approximation problems. Our approach i…