4 citations · 12 across the 13 of their papers we have counts for
3 papers · 1 filter
A view of mini-batch SGD via generating functions: conditions of convergence, phase transitions, benefit from negative momenta
Maksim Velikanov, Denis Kuznedelev, Dmitry Yarotsky
Mini-batch SGD with momentum is a fundamental algorithm for learning large predictive models. In this paper we develop a new analytic framework to analyze noise-averaged properties…
Embedded Ensembles: Infinite Width Limit and Operating Regimes
Maksim Velikanov, Roman Kail, Ivan Anokhin +4
A memory efficient approach to ensembling neural networks is to share most weights among the ensembled models by means of a single reference network. We refer to this strategy as E…
Tight Convergence Rate Bounds for Optimization Under Power Law Spectral Conditions
Maksim Velikanov, Dmitry Yarotsky
Performance of optimization on quadratic problems sensitively depends on the low-lying part of the spectrum. For large (effectively infinite-dimensional) problems, this part of the…