GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
arXiv:1706.08500
Abstract
Generative Adversarial Networks (GANs) excel at creating realistic images with complex models for which maximum likelihood is infeasible. However, the convergence of GAN training has still not been proved. We propose a two time-scale update rule (TTUR) for training GANs with stochastic gradient descent on arbitrary GAN loss functions. TTUR has an individual learning rate for both the discriminator and the generator. Using the theory of stochastic approximation, we prove that the TTUR converges under mild assumptions to a stationary local Nash equilibrium. The convergence carries over to the popular Adam optimization, for which we prove that it follows the dynamics of a heavy ball with friction and thus prefers flat minima in the objective landscape. For the evaluation of the performance of GANs at image generation, we introduce the "Fréchet Inception Distance" (FID) which captures the similarity of generated images to real ones better than the Inception Score. In experiments, TTUR improves learning for DCGANs and Improved Wasserstein GANs (WGAN-GP) outperforming conventional GAN training on CelebA, CIFAR-10, SVHN, LSUN Bedrooms, and the One Billion Word Benchmark.
Implementations are available at: https://github.com/bioinf-jku/TTUR
Cited by in corpus (275)
- Deep learning for molecular design - a review of the state of the art
- EdgeConnect: Generative Image Inpainting with Adversarial Edge Learning
- Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity
- Guided Image Generation with Conditional Invertible Neural Networks
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- GANSynth: Adversarial Neural Audio Synthesis
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Towards the Automatic Anime Characters Creation with Generative Adversarial Networks
- Contrastive Learning for Unpaired Image-to-Image Translation
- Freeze the Discriminator: a Simple Baseline for Fine-Tuning GANs
- Generating Diverse High-Fidelity Images with VQ-VAE-2
- High Fidelity Speech Synthesis with Adversarial Networks
- High-Fidelity Image Generation With Fewer Labels
- ShapeAssembly: Learning to Generate Programs for 3D Shape Structure Synthesis
- On Finding Local Nash Equilibria (and Only Local Nash Equilibria) in Zero-Sum Games
- Learning to Generate Diverse Dance Motions with Transformer
- SoftAdapt: Techniques for Adaptive Loss Weighting of Neural Networks with Multi-Part Loss Functions
- Sliced Score Matching: A Scalable Approach to Density and Score Estimation
- Conditional GAN for timeseries generation
- Maximum Entropy Generators for Energy-Based Models
- Compressing GANs using Knowledge Distillation
- Task Agnostic Continual Learning via Meta Learning
- Spatial Evolutionary Generative Adversarial Networks
- Privacy Protection in Street-View Panoramas using Depth and Multi-View Imagery
- Lightweight Modules for Efficient Deep Learning based Image Restoration
- Finger-GAN: Generating Realistic Fingerprint Images Using Connectivity Imposed GAN
- MeshGAN: Non-linear 3D Morphable Models of Faces
- Consistency Regularization for Generative Adversarial Networks
- Prescribed Generative Adversarial Networks
- Object-driven Text-to-Image Synthesis via Adversarial Training
- DM-GAN: Dynamic Memory Generative Adversarial Networks for Text-to-Image Synthesis
- DermGAN: Synthetic Generation of Clinical Skin Images with Pathology
- Semantic Hierarchy Emerges in Deep Generative Representations for Scene Synthesis
- Convolution with even-sized kernels and symmetric padding
- Lipschitz Generative Adversarial Nets
- Mode Seeking Generative Adversarial Networks for Diverse Image Synthesis
- Inverse Graphics GAN: Learning to Generate 3D Shapes from Unstructured 2D Data
- E-LPIPS: Robust Perceptual Image Similarity via Random Transformation Ensembles
- LiDAR Sensor modeling and Data augmentation with GANs for Autonomous driving
- Small-GAN: Speeding Up GAN Training Using Core-sets
- MelGAN-VC: Voice Conversion and Audio Style Transfer on arbitrarily long samples using Spectrograms
- Iris-GAN: Learning to Generate Realistic Iris Images Using Convolutional GAN
- Generating Multiple Objects at Spatially Distinct Locations
- Rethinking Image Inpainting via a Mutual Encoder-Decoder with Feature Equalizations
- Complement Face Forensic Detection and Localization with FacialLandmarks
- DeepFlow: History Matching in the Space of Deep Generative Models
- GANs May Have No Nash Equilibria
- GAN-QP: A Novel GAN Framework without Gradient Vanishing and Lipschitz Constraint
- Improving Inversion and Generation Diversity in StyleGAN using a Gaussianized Latent Space
- X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers
- Neural Language Generation: Formulation, Methods, and Evaluation
- FCC-GAN: A Fully Connected and Convolutional Net Architecture for GANs
- Augmented Normalizing Flows: Bridging the Gap Between Generative Flows and Latent Variable Models
- Hybrid Discriminative-Generative Training via Contrastive Learning
- Adversarial Learning for Improved Onsets and Frames Music Transcription
- Encoding Invariances in Deep Generative Models
- From Here to There: Video Inbetweening Using Direct 3D Convolutions
- Using Scene Graph Context to Improve Image Generation
- Smoothness and Stability in GANs
- High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks
- Fast and Provable ADMM for Learning with Generative Priors
- LaFIn: Generative Landmark Guided Face Inpainting
- Text and Style Conditioned GAN for Generation of Offline Handwriting Lines
- Exploring the Evolution of GANs through Quality Diversity
- Conservation AI: Live Stream Analysis for the Detection of Endangered Species Using Convolutional Neural Networks and Drone Technology
- On Solving Minimax Optimization Locally: A Follow-the-Ridge Approach
- Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
- Guided Image-to-Image Translation with Bi-Directional Feature Transformation
- Interactive Sketch & Fill: Multiclass Sketch-to-Image Translation
- Real or Not Real, that is the Question
- The Six Fronts of the Generative Adversarial Networks
- SRGAN: Training Dataset Matters
- A Utility-Preserving GAN for Face Obscuration
- GramGAN: Deep 3D Texture Synthesis From 2D Exemplars
- Improving the Speed and Quality of GAN by Adversarial Training
- Latent Translation: Crossing Modalities by Bridging Generative Models
- Video Generation from Single Semantic Label Map
- Unbalanced GANs: Pre-training the Generator of Generative Adversarial Network using Variational Autoencoder
- Example-Guided Style Consistent Image Synthesis from Semantic Labeling
- Explanation by Progressive Exaggeration
- AdvSPADE: Realistic Unrestricted Attacks for Semantic Segmentation
- Adversarial Generation of Time-Frequency Features with application in audio synthesis
- Unsupervised Medical Image Segmentation with Adversarial Networks: From Edge Diagrams to Segmentation Maps
- FTGAN: A Fully-trained Generative Adversarial Networks for Text to Face Generation
- Self-Adversarial Learning with Comparative Discrimination for Text Generation
- Jointly Measuring Diversity and Quality in Text Generation Models
- Deep Exemplar-based Video Colorization
- MulGAN: Facial Attribute Editing by Exemplar
- Free-form Video Inpainting with 3D Gated Convolution and Temporal PatchGAN
- Investigating Under and Overfitting in Wasserstein Generative Adversarial Networks
- Texture Mixer: A Network for Controllable Synthesis and Interpolation of Texture
- PI-REC: Progressive Image Reconstruction Network With Edge and Color Domain
- Pixel-wise Conditioned Generative Adversarial Networks for Image Synthesis and Completion
- DwNet: Dense warp-based network for pose-guided human video generation
- EncryptGAN: Image Steganography with Domain Transform
- Mask-Guided Portrait Editing with Conditional GANs
- An Improved Self-supervised GAN via Adversarial Training
- Towards Efficient and Unbiased Implementation of Lipschitz Continuity in GANs
- Wav2Pix: Speech-conditioned Face Generation using Generative Adversarial Networks
- GAN- vs. JPEG2000 Image Compression for Distributed Automotive Perception: Higher Peak SNR Does Not Mean Better Semantic Segmentation
- (q,p)-Wasserstein GANs: Comparing Ground Metrics for Wasserstein GANs
- Regularized Autoencoders via Relaxed Injective Probability Flow
- Controllable and Progressive Image Extrapolation
- A Survey and Taxonomy of Adversarial Neural Networks for Text-to-Image Synthesis
- Enhancing the Performance of Practical Profiling Side-Channel Attacks Using Conditional Generative Adversarial Networks
- Semantic Bottleneck Scene Generation
- RPGAN: GANs Interpretability via Random Routing
- Learning Spatial Pyramid Attentive Pooling in Image Synthesis and Image-to-Image Translation
- Learning latent representations across multiple data domains using Lifelong VAEGAN
- Neural Architecture Search for Deep Image Prior
- GAN Slimming: All-in-One GAN Compression by A Unified Optimization Framework
- Few-shot Video-to-Video Synthesis
- COCO-FUNIT: Few-Shot Unsupervised Image Translation with a Content Conditioned Style Encoder
- FLNet: Landmark Driven Fetching and Learning Network for Faithful Talking Facial Animation Synthesis
- Describe What to Change: A Text-guided Unsupervised Image-to-Image Translation Approach
- Invert and Defend: Model-based Approximate Inversion of Generative Adversarial Networks for Secure Inference
- On Data Augmentation and Adversarial Risk: An Empirical Analysis
- Super-resolution Variational Auto-Encoders
- Gradient Descent-Ascent Provably Converges to Strict Local Minmax Equilibria with a Finite Timescale Separation
- RankGAN: A Maximum Margin Ranking GAN for Generating Faces
- GAN-based Generation and Automatic Selection of Explanations for Neural Networks
- Transposer: Universal Texture Synthesis Using Feature Maps as Transposed Convolution Filter
- Adversarially Approximated Autoencoder for Image Generation and Manipulation
- Unpaired Image Super-Resolution using Pseudo-Supervision
- Causal Adversarial Network for Learning Conditional and Interventional Distributions
- TSIT: A Simple and Versatile Framework for Image-to-Image Translation
- Tdcgan: Temporal Dilated Convolutional Generative Adversarial Network for End-to-end Speech Enhancement
- MagGAN: High-Resolution Face Attribute Editing with Mask-Guided Generative Adversarial Network
- Rethinking Image Deraining via Rain Streaks and Vapors
- Privacy Leakage of SIFT Features via Deep Generative Model based Image Reconstruction
- Constrained Generative Adversarial Network Ensembles for Sharable Synthetic Data Generation
- Generative Adversarial Zero-Shot Relational Learning for Knowledge Graphs
- Conditional Adversarial Generative Flow for Controllable Image Synthesis
- Neural Turtle Graphics for Modeling City Road Layouts
- PMC-GANs: Generating Multi-Scale High-Quality Pedestrian with Multimodal Cascaded GANs
- Using Simulated Data to Generate Images of Climate Change
- Bidirectional Mapping Generative Adversarial Networks for Brain MR to PET Synthesis
- Task-agnostic Temporally Consistent Facial Video Editing
- Learning Neurosymbolic Generative Models via Program Synthesis
- Wasserstein-Wasserstein Auto-Encoders
- Natural and Realistic Single Image Super-Resolution with Explicit Natural Manifold Discrimination
- Biphasic Learning of GANs for High-Resolution Image-to-Image Translation
- BézierSketch: A generative model for scalable vector sketches
- Generated Loss and Augmented Training of MNIST VAE
- Reference-guided Face Component Editing
- Conditional WGANs with Adaptive Gradient Balancing for Sparse MRI Reconstruction
- Quality Evaluation of GANs Using Cross Local Intrinsic Dimensionality
- Coconditional Autoencoding Adversarial Networks for Chinese Font Feature Learning
- Learning Implicit Generative Models with Theoretical Guarantees
- Fast Universal Style Transfer for Artistic and Photorealistic Rendering
- A gradual, semi-discrete approach to generative network training via explicit Wasserstein minimization
- Learning to Incorporate Structure Knowledge for Image Inpainting
- Generative View Synthesis: From Single-view Semantics to Novel-view Images
- Selective Synthetic Augmentation with Quality Assurance
- The Art of Food: Meal Image Synthesis from Ingredients
- Dual Attention GANs for Semantic Image Synthesis
- Approximating Human Judgment of Generated Image Quality
- NAS-DIP: Learning Deep Image Prior with Neural Architecture Search
- Talking-head Generation with Rhythmic Head Motion
- RetrieveGAN: Image Synthesis via Differentiable Patch Retrieval
- Deep Generative Learning via Variational Gradient Flow
- PowerGAN: Synthesizing Appliance Power Signatures Using Generative Adversarial Networks
- Kernel Stein Generative Modeling
- Dual Generator Generative Adversarial Networks for Multi-Domain Image-to-Image Translation
- UVA: A Universal Variational Framework for Continuous Age Analysis
- StyleNAS: An Empirical Study of Neural Architecture Search to Uncover Surprisingly Fast End-to-End Universal Style Transfer Networks
- Adversarial Pixel-Level Generation of Semantic Images
- Learning Disentangled Representations with Latent Variation Predictability
- Evaluating Generative Adversarial Networks on Explicitly Parameterized Distributions
- Very Long Natural Scenery Image Prediction by Outpainting
- Conditional Denoising of Remote Sensing Imagery Using Cycle-Consistent Deep Generative Models
- Deep Fusion Network for Image Completion
- F2GAN: Fusing-and-Filling GAN for Few-shot Image Generation
- Domain-Specific Mappings for Generative Adversarial Style Transfer
- CAD-PU: A Curvature-Adaptive Deep Learning Solution for Point Set Upsampling
- Image-to-Image Translation with Multi-Path Consistency Regularization
- World-Consistent Video-to-Video Synthesis
- FaceShapeGene: A Disentangled Shape Representation for Flexible Face Image Editing
- Improving Style-Content Disentanglement in Image-to-Image Translation
- Example-Guided Scene Image Synthesis using Masked Spatial-Channel Attention and Patch-Based Self-Supervision
- Toward Joint Image Generation and Compression using Generative Adversarial Networks
- PriorGAN: Real Data Prior for Generative Adversarial Nets
- Multivariate-Information Adversarial Ensemble for Scalable Joint Distribution Matching
- Controllable Image Synthesis via SegVAE
- Exploring Generative Physics Models with Scientific Priors in Inertial Confinement Fusion
- Spherical Image Generation from a Single Normal Field of View Image by Considering Scene Symmetry
- Sample weighting as an explanation for mode collapse in generative adversarial networks
- Improving the Evaluation of Generative Models with Fuzzy Logic
- "Best-of-Many-Samples" Distribution Matching
- TinyGAN: Distilling BigGAN for Conditional Image Generation
- Orthogonal Wasserstein GANs
- Auto-Embedding Generative Adversarial Networks for High Resolution Image Synthesis
- Client Adaptation improves Federated Learning with Simulated Non-IID Clients
- Unsupervised Multi-Domain Multimodal Image-to-Image Translation with Explicit Domain-Constrained Disentanglement
- Optimal Transport Relaxations with Application to Wasserstein GANs
- Deconstructing Generative Adversarial Networks
- BreGMN: scaled-Bregman Generative Modeling Networks
- Unselfie: Translating Selfies to Neutral-pose Portraits in the Wild
- Searching for an (un)stable equilibrium: experiments in training generative models without data
- Intelligent Home 3D: Automatic 3D-House Design from Linguistic Descriptions Only
- Neural Crossbreed: Neural Based Image Metamorphosis
- CookGAN: Meal Image Synthesis from Ingredients
- An Empirical Study of Generative Models with Encoders
- Sampling Using Neural Networks for colorizing the grayscale images
- AE-OT-GAN: Training GANs from data specific latent distribution
- Face Attribute Invertion
- Characteristic Regularisation for Super-Resolving Face Images
- Unsupervised Adversarial Image Inpainting
- Detecting GAN generated errors
- A Layer-Based Sequential Framework for Scene Generation with GANs
- DepthwiseGANs: Fast Training Generative Adversarial Networks for Realistic Image Synthesis
- Robust Conditional GAN from Uncertainty-Aware Pairwise Comparisons
- Optimal Transport Based Generative Autoencoders
- A Generative Approach Towards Improved Robotic Detection of Marine Litter
- On the Anomalous Generalization of GANs
- One-Shot Image-to-Image Translation via Part-Global Learning with a Multi-adversarial Framework
- Kernel Mean Matching for Content Addressability of GANs
- Less Memory, Faster Speed: Refining Self-Attention Module for Image Reconstruction
- CariMe: Unpaired Caricature Generation with Multiple Exaggerations
- Generative Model without Prior Distribution Matching
- Improving Text to Image Generation using Mode-seeking Function
- Pairwise-GAN: Pose-based View Synthesis through Pair-Wise Training
- LaDDer: Latent Data Distribution Modelling with a Generative Prior
- Person-in-Context Synthesiswith Compositional Structural Space
- Channel-Directed Gradients for Optimization of Convolutional Neural Networks
- DeepLandscape: Adversarial Modeling of Landscape Video
- Encoder-Powered Generative Adversarial Networks
- DeepGIN: Deep Generative Inpainting Network for Extreme Image Inpainting
- Synthesizing 3D Shapes from Silhouette Image Collections using Multi-projection Generative Adversarial Networks
- Interpreting Spatially Infinite Generative Models
- Novel View Synthesis on Unpaired Data by Conditional Deformable Variational Auto-Encoder
- LEED: Label-Free Expression Editing via Disentanglement
- Benefiting Deep Latent Variable Models via Learning the Prior and Removing Latent Regularization
- Learning End-to-End Action Interaction by Paired-Embedding Data Augmentation
- Shapes and Context: In-the-Wild Image Synthesis & Manipulation
- Modeling Artistic Workflows for Image Generation and Editing
- Words as Art Materials: Generating Paintings with Sequential GANs
- Distributional Discrepancy: A Metric for Unconditional Text Generation
- Semi-supervised mp-MRI Data Synthesis with StitchLayer and Auxiliary Distance Maximization
- 3D Dense Geometry-Guided Facial Expression Synthesis by Adversarial Learning
- Improving Relational Regularized Autoencoders with Spherical Sliced Fused Gromov Wasserstein
- The Benefits of Pairwise Discriminators for Adversarial Training
- MCMI: Multi-Cycle Image Translation with Mutual Information Constraints
- Toward Learning a Unified Many-to-Many Mapping for Diverse Image Translation
- Conditional Transferring Features: Scaling GANs to Thousands of Classes with 30% Less High-quality Data for Training
- Generating Object Stamps
- Stacked Wasserstein Autoencoder
- Cross-Identity Motion Transfer for Arbitrary Objects through Pose-Attentive Video Reassembling
- Semantic Example Guided Image-to-Image Translation
- Comprehensive Facial Expression Synthesis using Human-Interpretable Language
- Neural Approximation of an Auto-Regressive Process through Confidence Guided Sampling
- Adversarial Code Learning for Image Generation
- LinesToFacePhoto: Face Photo Generation from Lines with Conditional Self-Attention Generative Adversarial Network
- A Wasserstein Minimum Velocity Approach to Learning Unnormalized Models
- Establishing an Evaluation Metric to Quantify Climate Change Image Realism
- Multimodal Image Outpainting With Regularized Normalized Diversification
- Learning to Inpaint by Progressively Growing the Mask Regions
- Joint Wasserstein Distribution Matching
- Wavelets to the Rescue: Improving Sample Quality of Latent Variable Deep Generative Models
- Pixel-wise Conditioning of Generative Adversarial Networks
- Human Annotations Improve GAN Performances
- Virtual Conditional Generative Adversarial Networks
- A novel generative reverse net assisted evolution algorithm for expensive-computational optimizations
- Improving Stability of LS-GANs for Audio and Speech Signals
- The Implicit Metropolis-Hastings Algorithm
- Semantic View Synthesis
- ByeGlassesGAN: Identity Preserving Eyeglasses Removal for Face Images
- One-element Batch Training by Moving Window
- Composition and decomposition of GANs
- Strategic Prediction with Latent Aggregative Games
- Video-to-Video Translation for Visual Speech Synthesis
- Translate the Facial Regions You Like Using Region-Wise Normalization
- Improving Generative Adversarial Networks with Local Coordinate Coding
- Toward Zero-Shot Unsupervised Image-to-Image Translation