Local AdaGrad-Type Algorithm for Stochastic Convex-Concave Optimization
arXiv:2106.10022
Abstract
Large scale convex-concave minimax problems arise in numerous applications, including game theory, robust training, and training of generative adversarial networks. Despite their wide applicability, solving such problems efficiently and effectively is challenging in the presence of large amounts of data using existing stochastic minimax methods. We study a class of stochastic minimax methods and develop a communication-efficient distributed stochastic extragradient algorithm, LocalAdaSEG, with an adaptive learning rate suitable for solving convex-concave minimax problems in the Parameter-Server model. LocalAdaSEG has three main features: (i) a periodic communication strategy that reduces the communication cost between workers and the server; (ii) an adaptive learning rate that is computed locally and allows for tuning-free implementation; and (iii) theoretically, a nearly linear speed-up with respect to the dominant variance term, arising from the estimation of the stochastic gradient, is proven in both the smooth and nonsmooth convex-concave settings. LocalAdaSEG is used to solve a stochastic bilinear game, and train a generative adversarial network. We compare LocalAdaSEG against several existing optimizers for minimax problems and demonstrate its efficacy through several experiments in both homogeneous and heterogeneous settings.
42 pages; Accepted to Machine Learning, 2022
References in corpus (11)
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- On the Convergence of Adam and Beyond
- Towards Principled Methods for Training Generative Adversarial Networks
- On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization
- A Universal Algorithm for Variational Inequalities Adaptive to Smoothness and Noise
- Distributionally Robust Federated Averaging
- A Single-Loop Smoothed Gradient Descent-Ascent Algorithm for Nonconvex-Concave Min-Max Problems
- Local Stochastic Gradient Descent Ascent: Convergence Analysis and Communication Efficiency
- Efficient Algorithms for Federated Saddle Point Optimization
- A Distributed Training Algorithm of Generative Adversarial Networks with Quantized Gradients
- Geometry-Aware Universal Mirror-Prox