1 paper · 1 filter
Christian Hirsch, Daniel Willhalm
We study large deviations in the context of stochastic gradient descent for one-hidden-layer neural networks with quadratic loss. We derive a quenched large deviation principle, wh…