Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Benchmarking Neural Network Training Algorithms
George E. Dahl, Frank Schneider, Zachary Nado +22
Training algorithms, broadly construed, are an essential part of every deep learning pipeline. Training algorithm improvements that speed up training across a wide variety of workl…
cs.LG2024
Adaptive Gradient Methods at the Edge of Stability
Jeremy M. Cohen, Behrooz Ghorbani, Shankar Krishnan +8
Very little is known about the training dynamics of adaptive gradient methods like Adam in deep learning. In this paper, we shed light on the behavior of these algorithms in the fu…