An optimal randomized incremental gradient method
arXiv:1507.02000
Abstract
In this paper, we consider a class of finite-sum convex optimization problems whose objective function is given by the summation of () smooth components together with some other relatively simple terms. We first introduce a deterministic primal-dual gradient (PDG) method that can achieve the optimal black-box iteration complexity for solving these composite optimization problems using a primal-dual termination criterion. Our major contribution is to develop a randomized primal-dual gradient (RPDG) method, which needs to compute the gradient of only one randomly selected smooth component at each iteration, but can possibly achieve better complexity than PDG in terms of the total number of gradient evaluations. More specifically, we show that the total number of gradient evaluations performed by RPDG can be times smaller, both in expectation and with high probability, than those performed by deterministic optimal first-order methods under favorable situations. We also show that the complexity of the RPDG method is not improvable by developing a new lower complexity bound for a general class of randomized methods for solving large-scale finite-sum convex optimization problems. Moreover, through the development of PDG and RPDG, we introduce a novel game-theoretic interpretation for these optimal methods for convex optimization.
References in corpus (2)
Cited by in corpus (24)
- Stochastic Variance Reduction for Nonconvex Optimization
- A Universal Catalyst for First-Order Optimization
- A Simple Proximal Stochastic Gradient Method for Nonsmooth Nonconvex Optimization
- Tight Complexity Bounds for Optimizing Composite Objectives
- Distributed Stochastic Variance Reduced Gradient Methods and A Lower Bound for Communication Complexity
- NESTT: A Nonconvex Primal-Dual Splitting Method for Distributed and Stochastic Optimization
- The Practicality of Stochastic Optimization in Imaging Inverse Problems
- The proximal point method revisited
- Proximal-Proximal-Gradient Method
- A Unified Analysis of Stochastic Gradient Methods for Nonconvex Federated Optimization
- PAGE: A Simple and Optimal Probabilistic Gradient Estimator for Nonconvex Optimization
- Exploiting Strong Convexity from Data with Primal-Dual First-Order Algorithms
- Communication-Efficient Algorithms for Decentralized and Stochastic Optimization
- ANITA: An Optimal Loopless Accelerated Variance-Reduced Gradient Method
- CANITA: Faster Rates for Distributed Convex Optimization with Communication Compression
- Sketching Meets Random Projection in the Dual: A Provable Recovery Algorithm for Big and High-dimensional Data
- Nestrov's Acceleration For Second Order Method
- Dimension-Free Iteration Complexity of Finite Sum Optimization Problems
- Nesterov's Acceleration For Approximate Newton
- Limitations on Variance-Reduction and Acceleration Schemes for Finite Sum Optimization
- Larger is Better: The Effect of Learning Rates Enjoyed by Stochastic Optimization with Progressive Variance Reduction
- Structure-Adaptive, Variance-Reduced, and Accelerated Stochastic Optimization
- Random gradient extrapolation for distributed and stochastic optimization
- Fast Global Convergence via Landscape of Empirical Loss