A Decentralized Parallel Algorithm for Training Generative Adversarial Nets
arXiv:1910.12999
Abstract
Generative Adversarial Networks (GANs) are a powerful class of generative models in the deep learning community. Current practice on large-scale GAN training utilizes large models and distributed large-batch training strategies, and is implemented on deep learning frameworks (e.g., TensorFlow, PyTorch, etc.) designed in a centralized manner. In the centralized network topology, every worker needs to either directly communicate with the central node or indirectly communicate with all other workers in every iteration. However, when the network bandwidth is low or network latency is high, the performance would be significantly degraded. Despite recent progress on decentralized algorithms for training deep neural networks, it remains unclear whether it is possible to train GANs in a decentralized manner. The main difficulty lies at handling the nonconvex-nonconcave min-max optimization and the decentralized communication simultaneously. In this paper, we address this difficulty by designing the \textbf{first gradient-based decentralized parallel algorithm} which allows workers to have multiple rounds of communications in one iteration and to update the discriminator and generator simultaneously, and this design makes it amenable for the convergence analysis of the proposed decentralized algorithm. Theoretically, our proposed decentralized algorithm is able to solve a class of non-convex non-concave min-max problems with provable non-asymptotic convergence to first-order stationary point. Experimental results on GANs demonstrate the effectiveness of the proposed algorithm.
Accepted by NeurIPS 2020
References in corpus (35)
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Large Scale GAN Training for High Fidelity Natural Image Synthesis
- Spectral Normalization for Generative Adversarial Networks
- Self-Attention Generative Adversarial Networks
- On the Linear Convergence of the ADMM in Decentralized Consensus Optimization
- Certifying Some Distributional Robustness with Principled Adversarial Training
- Adversarial Learning for Neural Dialogue Generation
- Optimal algorithms for smooth and strongly convex distributed optimization in networks
- Decentralized Stochastic Optimization and Gossip Algorithms with Compressed Communication
- On Gradient Descent Ascent for Nonconvex-Concave Minimax Problems
- D: Decentralized Training over Decentralized Data
- Cooperative SGD: A unified Framework for the Design and Analysis of Communication-Efficient SGD Algorithms
- A Variational Inequality Perspective on Generative Adversarial Networks
- DSA: Decentralized Double Stochastic Averaging Gradient Algorithm
- On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization
- Optimization, Learning, and Games with Predictable Sequences
- Multi-Agent Reinforcement Learning via Double Averaging Primal-Dual Optimization
- Optimistic mirror descent in saddle-point problems: Going the extra (gradient) mile
- TAC-GAN - Text Conditioned Auxiliary Classifier Generative Adversarial Network
- On Finding Local Nash Equilibria (and Only Local Nash Equilibria) in Zero-Sum Games
- Reducing Noise in GAN Training with Variance Reduced Extragradient
- Competitive Gradient Descent
- MATCHA: Speeding Up Decentralized SGD via Matching Decomposition Sampling
- An Alternative View: When Does SGD Escape Local Minima?
- Decentralized Deep Learning with Arbitrary Communication Compression
- Efficient Algorithms for Smooth Minimax Optimization
- SGD Converges to Global Minimum in Deep Learning via Star-convex Path
- Towards Better Understanding of Adaptive Gradient Algorithms in Generative Adversarial Nets
- Accelerated Decentralized Optimization with Local Updates for Smooth and Strongly Convex Objectives
- A Decentralized Proximal Point-type Method for Saddle Point Problems
- An Online Learning Approach to Generative Adversarial Networks
- Solving Non-Convex Non-Concave Min-Max Games Under Polyak-Łojasiewicz Condition
- SPARQ-SGD: Event-Triggered and Compressed Communication in Decentralized Stochastic Optimization
- Variance-Reduced Decentralized Stochastic Optimization with Gradient Tracking--Part I: GT-SAGA
- Distributed Deep Learning Strategies For Automatic Speech Recognition
Cited by in corpus (6)
- Recent theoretical advances in decentralized distributed convex optimization
- Local Stochastic Gradient Descent Ascent: Convergence Analysis and Communication Efficiency
- Distributed Saddle-Point Problems Under Similarity
- A Decentralized Adaptive Momentum Method for Solving a Class of Min-Max Optimization Problems
- Near-Optimal Decentralized Algorithms for Saddle Point Problems over Time-Varying Networks
- CDMA: A Practical Cross-Device Federated Learning Algorithm for General Minimax Problems