3 citations · 8 across the 16 of their papers we have counts for
3 papers · 1 filter
DAG: Projected Stochastic Approximation Iteration for DAG Structure Learning
Klea Ziu, Slavomír Hanzely, Loka Li +3
Learning the structure of Directed Acyclic Graphs (DAGs) presents a significant challenge due to the vast combinatorial search space of possible graphs, which scales exponentially…
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
Petr Ostroukhov, Aigerim Zhumabayeva, Chulu Xiang +3
This paper presents a novel adaptation of the Stochastic Gradient Descent (SGD), termed AdaBatchGrad. This modification seamlessly integrates an adaptive step size with an adjustab…
SANIA: Polyak-type Optimization Framework Leads to Scale Invariant Stochastic Algorithms
Farshed Abdukhakimov, Chulu Xiang, Dmitry Kamzolov +2
Adaptive optimization methods are widely recognized as among the most popular approaches for training Deep Neural Networks (DNNs). Techniques such as Adam, AdaGrad, and AdaHessian…