1 paper
Ioannis Bantzis, James B. Simon, Arthur Jacot
When a deep ReLU network is initialized with small weights, gradient descent (GD) is at first dominated by the saddle at the origin in parameter space. We study the so-called escap…