2 citations · 2 across the 5 of their papers we have counts for
5 papers
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
Petr Ostroukhov, Aigerim Zhumabayeva, Chulu Xiang +3
This paper presents a novel adaptation of the Stochastic Gradient Descent (SGD), termed AdaBatchGrad. This modification seamlessly integrates an adaptive step size with an adjustab…
SANIA: Polyak-type Optimization Framework Leads to Scale Invariant Stochastic Algorithms
Farshed Abdukhakimov, Chulu Xiang, Dmitry Kamzolov +2
Adaptive optimization methods are widely recognized as among the most popular approaches for training Deep Neural Networks (DNNs). Techniques such as Adam, AdaGrad, and AdaHessian…
Stochastic Gradient Descent with Preconditioned Polyak Step-size
Farshed Abdukhakimov, Chulu Xiang, Dmitry Kamzolov +1
Stochastic Gradient Descent (SGD) is one of the many iterative optimization methods that are widely used in solving machine learning problems. These methods display valuable proper…
Cubic Regularization is the Key! The First Accelerated Quasi-Newton Method with a Global Convergence Rate of for Convex Functions
Dmitry Kamzolov, Klea Ziu, Artem Agafonov +1
In this paper, we propose the first Quasi-Newton method with a global convergence rate of for general convex functions. Quasi-Newton methods, such as BFGS, SR-1, are we…
Suppressing Poisoning Attacks on Federated Learning for Medical Imaging
Naif Alkhunaizi, Dmitry Kamzolov, Martin Takáč +1
Collaboration among multiple data-owning entities (e.g., hospitals) can accelerate the training process and yield better machine learning models due to the availability and diversi…