3 papers
cs.LG2026
Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics
Igor Ignashin, Anna Radovskaya, Andrew Semenov +7
Stochastic Gradient Descent (SGD) is commonly modeled as a Langevin process, assuming that minibatch noise acts as Brownian motion. However, this approximation relies on a continuo…
math.OC2025
Adaptive Regularized Newton Method with Inexact Hessian
Aleksandr Shestakov, Nail Bashirov, Andrei Semenov +4
Newton's method is the most widespread high-order method, demanding the gradient and the Hessian of the objective function. However, one of the main disadvantages of Newtons method…
math.OC2025
Unified Theory of Adaptive Variance Reduction
Aleksandr Shestakov, Valery Parfenov, Aleksandr Beznosikov
Variance reduction is a family of powerful mechanisms for stochastic optimization that appears to be helpful in many machine learning tasks. It is based on estimating the exact gra…