The Cramer Distance as a Solution to Biased Wasserstein Gradients
arXiv:1705.10743
Abstract
The Wasserstein probability metric has received much attention from the machine learning community. Unlike the Kullback-Leibler divergence, which strictly measures change in probability, the Wasserstein metric reflects the underlying geometry between outcomes. The value of being sensitive to this geometry has been demonstrated, among others, in ordinal regression and generative modelling. In this paper we describe three natural properties of probability divergences that reflect requirements from machine learning: sum invariance, scale sensitivity, and unbiased sample gradients. The Wasserstein metric possesses the first two properties but, unlike the Kullback-Leibler divergence, does not possess the third. We provide empirical evidence suggesting that this is a serious issue in practice. Leveraging insights from probabilistic forecasting we propose an alternative to the Wasserstein metric, the Cramér distance. We show that the Cramér distance possesses all three desired properties, combining the best of the Wasserstein and Kullback-Leibler divergences. To illustrate the relevance of the Cramér distance in practice we design a new algorithm, the Cramér Generative Adversarial Network (GAN), and show that it performs significantly better than the related Wasserstein GAN.
References in corpus (1)
Cited by in corpus (34)
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- A Distributional Perspective on Reinforcement Learning
- Towards the Automatic Anime Characters Creation with Generative Adversarial Networks
- ChemGAN challenge for drug discovery: can AI reproduce natural chemical diversity?
- DeepMDP: Learning Continuous Latent Space Models for Representation Learning
- Lipschitz Generative Adversarial Nets
- Unbalanced minibatch Optimal Transport; applications to Domain Adaptation
- Small-GAN: Speeding Up GAN Training Using Core-sets
- Cherenkov Detectors Fast Simulation Using Neural Networks
- Learning Generative Models across Incomparable Spaces
- A Comparative Analysis of Expected and Distributional Reinforcement Learning
- A Survey on Optimal Transport for Machine Learning: Theory and Applications
- Minibatch optimal transport distances; analysis and applications
- Towards Robust, Locally Linear Deep Networks
- Disentangled Recurrent Wasserstein Autoencoder
- The Six Fronts of the Generative Adversarial Networks
- ChildGAN: Large Scale Synthetic Child Facial Data Using Domain Adaptation in StyleGAN
- On Wasserstein Reinforcement Learning and the Fokker-Planck equation
- How Do the Hearts of Deep Fakes Beat? Deep Fake Source Detection via Interpreting Residuals with Biological Signals
- Do Neural Optimal Transport Solvers Work? A Continuous Wasserstein-2 Benchmark
- Likelihood Estimation for Generative Adversarial Networks
- Estimating and Evaluating Regression Predictive Uncertainty in Deep Object Detectors
- Metric Learning-based Generative Adversarial Network
- FakePolisher: Making DeepFakes More Detection-Evasive by Shallow Reconstruction
- A gradual, semi-discrete approach to generative network training via explicit Wasserstein minimization
- Generalized Sliced Distances for Probability Distributions
- KALE Flow: A Relaxed KL Gradient Flow for Probabilities with Disjoint Support
- Pareto GAN: Extending the Representational Power of GANs to Heavy-Tailed Distributions
- Alignment Attention by Matching Key and Query Distributions
- PC-GAIN: Pseudo-label Conditional Generative Adversarial Imputation Networks for Incomplete Data
- RSD-GAN: Regularized Sobolev Defense GAN Against Speech-to-Text Adversarial Attacks
- Implicit Manifold Learning on Generative Adversarial Networks
- Complexity Controlled Generative Adversarial Networks
- SafeML: Safety Monitoring of Machine Learning Classifiers through Statistical Difference Measure