A Decentralized Adaptive Momentum Method for Solving a Class of Min-Max Optimization Problems
arXiv:2106.06075 · doi:10.1016/j.sigpro.2021.108245
Abstract
Min-max saddle point games have recently been intensely studied, due to their wide range of applications, including training Generative Adversarial Networks (GANs). However, most of the recent efforts for solving them are limited to special regimes such as convex-concave games. Further, it is customarily assumed that the underlying optimization problem is solved either by a single machine or in the case of multiple machines connected in centralized fashion, wherein each one communicates with a central node. The latter approach becomes challenging, when the underlying communications network has low bandwidth. In addition, privacy considerations may dictate that certain nodes can communicate with a subset of other nodes. Hence, it is of interest to develop methods that solve min-max games in a decentralized manner. To that end, we develop a decentralized adaptive momentum (ADAM)-type algorithm for solving min-max optimization problem under the condition that the objective function satisfies a Minty Variational Inequality condition, which is a generalization to convex-concave case. The proposed method overcomes shortcomings of recent non-adaptive gradient-based decentralized algorithms for min-max optimization problems that do not perform well in practice and require careful tuning. In this paper, we obtain non-asymptotic rates of convergence of the proposed algorithm (coined DADAM) for finding a (stochastic) first-order Nash equilibrium point and subsequently evaluate its performance on training GANs. The extensive empirical evaluation shows that DADAM outperforms recently developed methods, including decentralized optimistic stochastic gradient for solving such min-max problems.
References in corpus (13)
- ADADELTA: An Adaptive Learning Rate Method
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- On the Convergence of Adam and Beyond
- Adaptive Bound Optimization for Online Convex Optimization
- Global Convergence and Variance-Reduced Optimization for a Class of Nonconvex-Nonconcave Minimax Problems
- GANs May Have No Nash Equilibria
- SGD Converges to Global Minimum in Deep Learning via Star-convex Path
- Frank-Wolfe Algorithms for Saddle Point Problems
- DADAM: A Consensus-based Distributed Adaptive Gradient Method for Online Optimization
- UniXGrad: A Universal, Adaptive Algorithm with Optimal Guarantees for Constrained Optimization
- A Decentralized Proximal Point-type Method for Saddle Point Problems
- Adaptive First-and Zeroth-order Methods for Weakly Convex Stochastic Optimization Problems
- Improving Efficiency in Large-Scale Decentralized Distributed Training