9 citations · 14 across the 15 of their papers we have counts for
Showing stat.MLShow all
2 papers · 1 filter
stat.ML2024
Emergence of heavy tails in homogenized stochastic gradient descent
Zhe Jiao, Martin Keller-Ressel
It has repeatedly been observed that loss minimization by stochastic gradient descent (SGD) leads to heavy-tailed distributions of neural network parameters. Here, we analyze a con…
stat.ML2020
A Theory of Hyperbolic Prototype Learning
Martin Keller-Ressel
We introduce Hyperbolic Prototype Learning, a type of supervised learning, where class labels are represented by ideal points (points at infinity) in hyperbolic space. Learning is…