On the Discrimination-Generalization Tradeoff in GANs
arXiv:1711.02771
Abstract
Generative adversarial training can be generally understood as minimizing certain moment matching loss defined by a set of discriminator functions, typically neural networks. The discriminator set should be large enough to be able to uniquely identify the true distribution (discriminative), and also be small enough to go beyond memorizing samples (generalizable). In this paper, we show that a discriminator set is guaranteed to be discriminative whenever its linear span is dense in the set of bounded continuous functions. This is a very mild condition satisfied even by neural networks with a single neuron. Further, we develop generalization bounds between the learned distribution and true distribution under different evaluation metrics. When evaluated with neural distance, our bounds show that generalization is guaranteed as long as the discriminator set is small enough, regardless of the size of the generator or hypothesis set. When evaluated with KL divergence, our bound provides an explanation on the counter-intuitive behaviors of testing likelihood in GAN training. Our analysis sheds lights on understanding the practical performance of GANs.
ICLR 2018
Cited by in corpus (25)
- Convergence Problems with Generative Adversarial Networks (GANs)
- Approximability of Discriminators Implies Diversity in GANs
- Adversarial Learning with Local Coordinate Coding
- On Catastrophic Forgetting and Mode Collapse in Generative Adversarial Networks
- 2-Wasserstein Approximation via Restricted Convex Potentials with Application to Improved Training for GANs
- Error Bounds of Imitating Policies and Environments
- A likelihood approach to nonparametric estimation of a singular distribution using deep generative models
- SGD Learns One-Layer Networks in WGANs
- Weighted Meta-Learning
- Limit Distribution Theory for the Smooth 1-Wasserstein Distance with Applications
- Smooth -Wasserstein Distance: Structure, Empirical Approximation, and Statistical Applications
- Asymptotic Guarantees for Generative Modeling Based on the Smooth Wasserstein Distance
- Train simultaneously, generalize better: Stability of gradient-based minimax learners
- Max-Affine Spline Insights into Deep Generative Networks
- Novel Human-Object Interaction Detection via Adversarial Domain Generalization
- Deep Dimension Reduction for Supervised Representation Learning
- KALE Flow: A Relaxed KL Gradient Flow for Probabilities with Disjoint Support
- RecurJac: An Efficient Recursive Algorithm for Bounding Jacobian Matrix of Neural Networks and Its Applications
- On Computation and Generalization of GANs with Spectrum Control
- On Value Discrepancy of Imitation Learning
- Distributional Robustness with IPMs and links to Regularization and GANs
- Semi-Implicit Generative Model
- Semantically Robust Unpaired Image Translation for Data with Unmatched Semantics Statistics
- Non-Asymptotic Error Bounds for Bidirectional GANs
- Neural Estimation of Statistical Divergences