2 papers
cs.LG2025
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
Ildus Sadrtdinov, Ivan Klimov, Ekaterina Lobacheva +1
We present a thermodynamic interpretation of the stationary behavior of stochastic gradient descent (SGD) under fixed learning rates (LRs) in neural network training. We show that…
cs.LG2024
Where Do Large Learning Rates Lead Us?
Ildus Sadrtdinov, Maxim Kodryan, Eduard Pokonechny +2
It is generally accepted that starting neural networks training with large learning rates (LRs) improves generalization. Following a line of research devoted to understanding this…