Adaptivity via a Parallel Architecture for Stochastic Gradient Methods
arXiv:2607.28902
Abstract
We develop a parallel framework that assembles static gradient methods to achieve better adaptivity. A static gradient method, denoted by , takes as input an initial point and specifying the number $\floor{T}$ of iterations. The step size is chosen as , where is a predetermined function of . The method then performs the iterations where is a stochastic gradient evaluated at , and is a scaling factor. For an integer , the processors in the proposed parallel framework search for an appropriate value of according to a geometric sequence so that the resulting gradient descent satisfies the desired convergence conditions. Each processor executes an infinite sequence of stages indexed by . At stage , processor is assigned where is a prescribed function. Processor executes at stage .