48 citations · 100 across the 9 of their papers we have counts for
3 papers · 1 filter
A Differentiable Rank-Based Objective For Better Feature Learning
Krunoslav Lehman Pavasovic, David Lopez-Paz, Giulio Biroli +1
In this paper, we leverage existing statistical methods to better understand feature learning from data. We tackle this by modifying the model-free variable selection method, Featu…
On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks
Umut Şimşekli, Mert Gürbüzbalaban, Thanh Huy Nguyen +2
The gradient noise (GN) in the stochastic gradient descent (SGD) algorithm is often considered to be Gaussian in the large data regime by assuming that the \emph{classical} central…
Comparing Dynamics: Deep Neural Networks versus Glassy Systems
M. Baity-Jesi, L. Sagun, M. Geiger +6
We analyze numerically the training dynamics of deep neural networks (DNN) by using methods developed in statistical physics of glassy systems. The two main issues we address are (…