Gradient penalty from a maximum margin perspective
arXiv:1910.06922
Abstract
A popular heuristic for improved performance in Generative adversarial networks (GANs) is to use some form of gradient penalty on the discriminator. This gradient penalty was originally motivated by a Wasserstein distance formulation. However, the use of gradient penalty in other GAN formulations is not well motivated. We present a unifying framework of expected margin maximization and show that a wide range of gradient-penalized GANs (e.g., Wasserstein, Standard, Least-Squares, and Hinge GANs) can be derived from this framework. Our results imply that employing gradient penalties induces a large-margin classifier (thus, a large-margin discriminator in GANs). We describe how expected margin maximization helps reduce vanishing gradients at fake (generated) samples, a known problem in GANs. From this framework, we derive a new gradient norm penalty with Hinge loss which generally produces equally good (or better) generated output in GANs than -norm penalties (based on the Fréchet Inception Distance).
Code at https://github.com/AlexiaJM/MaximumMarginGANs
References in corpus (8)
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Large Scale GAN Training for High Fidelity Natural Image Synthesis
- Spectral Normalization for Generative Adversarial Networks
- Which Training Methods for GANs do actually Converge?
- The relativistic discriminator: a key element missing from standard GAN
- On the regularization of Wasserstein GANs
- Banach Wasserstein GAN
- On Relativistic -Divergences