3 papers
cs.LG2026
Learning Rate Transfer in Normalized Transformers
Boris Shigida, Boris Hanin, Andrey Gromov
The Normalized Transformer, or nGPT (arXiv:2410.01131) achieves impressive training speedups and does not require weight decay or learning rate warmup. However, despite having hype…
cs.LG2025
Modified Loss of Momentum Gradient Descent: Fine-Grained Analysis
Matias D. Cattaneo, Boris Shigida
We analyze gradient descent with Polyak heavy-ball momentum (HB) whose fixed momentum parameter provides exponential decay of memory. Building on Kovachki and Stuart…
cs.LG2025
How Memory in Optimization Algorithms Implicitly Modifies the Loss
Matias D. Cattaneo, Boris Shigida
In modern optimization methods used in deep learning, each update depends on the history of previous iterations, often referred to as memory, and this dependence decays fast as the…