4 papers
Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
Ildus Sadrtdinov, Ekaterina Lobacheva, Ivan Klimov +3
Understanding the training dynamics of deep neural networks remains a major open problem, with physics-inspired approaches offering promising insights. Building on this perspective…
Why Gaussian Diffusion Models Fail on Discrete Data and How to Prevent It?
Alexander Shabalin, Simon Elistratov, Viacheslav Meshchaninov +2
Diffusion models have become a standard approach for generative modeling in continuous domains, yet their application to discrete data remains challenging. We investigate why Gauss…
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
Ildus Sadrtdinov, Ivan Klimov, Ekaterina Lobacheva +1
We present a thermodynamic interpretation of the stationary behavior of stochastic gradient descent (SGD) under fixed learning rates (LRs) in neural network training. We show that…
Where Do Large Learning Rates Lead Us?
Ildus Sadrtdinov, Maxim Kodryan, Eduard Pokonechny +2
It is generally accepted that starting neural networks training with large learning rates (LRs) improves generalization. Following a line of research devoted to understanding this…