Training Deep Energy-Based Models with f-Divergence Minimization
arXiv:2003.03463
Abstract
Deep energy-based models (EBMs) are very flexible in distribution parametrization but computationally challenging because of the intractable partition function. They are typically trained via maximum likelihood, using contrastive divergence to approximate the gradient of the KL divergence between data and model distribution. While KL divergence has many desirable properties, other f-divergences have shown advantages in training implicit density generative models such as generative adversarial networks. In this paper, we propose a general variational framework termed f-EBM to train EBMs using any desired f-divergence. We introduce a corresponding optimization algorithm and prove its local convergence property with non-linear dynamical systems theory. Experimental results demonstrate the superiority of f-EBM over contrastive divergence, as well as the benefits of training EBMs using f-divergences other than KL.
ICML 2020
References in corpus (7)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- WaveNet: A Generative Model for Raw Audio
- MADE: Masked Autoencoder for Distribution Estimation
- Learning in Implicit Generative Models
- Learning to Draw Samples: With Application to Amortized MLE for Generative Adversarial Learning
- Maximum Entropy Generators for Energy-Based Models
- Noise contrastive estimation: asymptotics, comparison with MC-MLE
Cited by in corpus (9)
- How to Train Your Energy-Based Models
- Improved Contrastive Divergence Training of Energy Based Models
- -GAIL: Learning -Divergence for Generative Adversarial Imitation Learning
- Telescoping Density-Ratio Estimation
- Generalized Energy Based Models
- Autoregressive Score Matching
- Learning Discrete Energy-based Models via Auxiliary-variable Local Exploration
- Imitation with Neural Density Models
- Analyzing and Improving the Optimization Landscape of Noise-Contrastive Estimation