activity
20192023
most citedA proof of convergence for gradient descent in the training of artificial neural networks for constant target functions

26 citations · 40 across the 4 of their papers we have counts for

collaborators

6 papers

math.NA2023

Efficient Sobolev approximation of linear parabolic PDEs in high dimensions

Patrick Cheridito, Florian Rossmannek

In this paper, we study the error in first order Sobolev norm in the approximation of solutions to linear parabolic PDEs. We use a Monte Carlo Euler scheme obtained from combining…

cs.LG2022★ 3 cited

Gradient descent provably escapes saddle points in the training of shallow ReLU networks

Patrick Cheridito, Arnulf Jentzen, Florian Rossmannek

Dynamical systems theory has recently been applied in optimization to prove that gradient descent algorithms bypass so-called strict saddle points of the loss function. However, in…

cs.LG2021★ 11 cited

Landscape analysis for shallow neural networks: complete classification of critical points for affine target functions

Patrick Cheridito, Arnulf Jentzen, Florian Rossmannek

In this paper, we analyze the landscape of the true loss of neural networks with one hidden layer and ReLU, leaky ReLU, or quadratic activation. In all three cases, we provide a co…

math.NA2021★ 26 cited

A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions

Patrick Cheridito, Arnulf Jentzen, Adrian Riekert +1

Gradient descent optimization algorithms are the standard ingredients that are used to train artificial neural networks (ANNs). Even though a huge number of numerical simulations i…

cs.LG2020

Non-convergence of stochastic gradient descent in the training of deep neural networks

Patrick Cheridito, Arnulf Jentzen, Florian Rossmannek

Deep neural networks have successfully been trained in various application areas with stochastic gradient descent. However, there exists no rigorous mathematical explanation why th…

math.NA2019

Efficient approximation of high-dimensional functions with neural networks

Patrick Cheridito, Arnulf Jentzen, Florian Rossmannek

In this paper, we develop a framework for showing that neural networks can overcome the curse of dimensionality in different high-dimensional approximation problems. Our approach i…