2 papers
stat.ML2024
Emergence of heavy tails in homogenized stochastic gradient descent
Zhe Jiao, Martin Keller-Ressel
It has repeatedly been observed that loss minimization by stochastic gradient descent (SGD) leads to heavy-tailed distributions of neural network parameters. Here, we analyze a con…
math.PR2021
Averaging principle of stochastic Burgers equation driven by Lévy processes
Hongge Yue, Yong Xu, Ruifang Wang +1
We are concerned about the averaging principle for the stochastic Burgers equation with slow-fast time scale. This slow-fast system is driven by Lévy processes. Under some appropri…