1 paper
Zhe Jiao, Martin Keller-Ressel
It has repeatedly been observed that loss minimization by stochastic gradient descent (SGD) leads to heavy-tailed distributions of neural network parameters. Here, we analyze a con…