2 papers
cs.LG2024
SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network
Tomer Galanti, Zachary S. Siegel, Aparna Gupte +1
We investigate the inherent bias of Stochastic Gradient Descent (SGD) toward learning low-rank weight matrices during the training of deep neural networks. Our results demonstrate…
stat.ML2024
Iterative regularization in classification via hinge loss diagonal descent
Vassilis Apidopoulos, Tomaso Poggio, Lorenzo Rosasco +1
Iterative regularization is a classic idea in regularization theory, that has recently become popular in machine learning. On the one hand, it allows to design efficient algorithms…