2 papers
cs.LG2024
MoMo: Momentum Models for Adaptive Learning Rates
Fabian Schaipp, Ruben Ohana, Michael Eickenberg +2
Training a modern machine learning architecture on a new task requires extensive learning-rate tuning, which comes at a high computational cost. Here we develop new Polyak-type ada…
cs.LG2024
Improving Convergence and Generalization Using Parameter Symmetries
Bo Zhao, Robert M. Gower, Robin Walters +1
In many neural networks, different values of the parameters may result in the same loss value. Parameter space symmetries are loss-invariant transformations that change the model p…