Skill Rating for Generative Models
arXiv:1808.04888
Abstract
We explore a new way to evaluate generative models using insights from evaluation of competitive games between human players. We show experimentally that tournaments between generators and discriminators provide an effective way to evaluate generative models. We introduce two methods for summarizing tournament outcomes: tournament win rate and skill rating. Evaluations are useful in different contexts, including monitoring the progress of a single model as it learns during the training process, and comparing the capabilities of two different fully trained models. We show that a tournament consisting of a single model playing against past and future versions of itself produces a useful measure of training progress. A tournament containing multiple separate models (using different seeds, hyperparameters, and architectures) provides a useful relative comparison between different trained GANs. Tournament-based rating methods are conceptually distinct from numerous previous categories of approaches to evaluation of generative models, and have complementary advantages and disadvantages.
References in corpus (9)
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- Progressive Growing of GANs for Improved Quality, Stability, and Variation
- Improved Techniques for Training GANs
- A Neural Representation of Sketch Drawings
- A note on the evaluation of generative models
- On distinguishability criteria for estimating generative models
- An Online Learning Approach to Generative Adversarial Networks
Cited by in corpus (12)
- Unifying Human and Statistical Evaluation for Natural Language Generation
- Small-GAN: Speeding Up GAN Training Using Core-sets
- Sketch-based Creativity Support Tools using Deep Learning
- Visual Intelligence through Human Interaction
- Improved Consistency Regularization for GANs
- Top-k Training of GANs: Improving GAN Performance by Throwing Away Bad Samples
- On Characterizing GAN Convergence Through Proximal Duality Gap
- Evolving SimGANs to Improve Abnormal Electrocardiogram Classification
- Manifold Topology Divergence: a Framework for Comparing Data Manifolds
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation Models
- Lipschitz Constrained GANs via Boundedness and Continuity
- Generative Models for Security: Attacks, Defenses, and Opportunities