Kernel Exponential Family Estimation via Doubly Dual Embedding
arXiv:1811.02228
Abstract
We investigate penalized maximum log-likelihood estimation for exponential family distributions whose natural parameter resides in a reproducing kernel Hilbert space. Key to our approach is a novel technique, doubly dual embedding, that avoids computation of the partition function. This technique also allows the development of a flexible sampling strategy that amortizes the cost of Monte-Carlo sampling in the inference stage. The resulting estimator can be easily generalized to kernel conditional exponential families. We establish a connection between kernel exponential family estimation and MMD-GANs, revealing a new perspective for understanding GANs. Compared to the score matching based estimators, the proposed method improves both memory and time efficiency while enjoying stronger statistical properties, such as fully capturing smoothness in its statistical convergence rate while the score matching estimator appears to saturate. Finally, we show that the proposed estimator empirically outperforms state-of-the-art
22 pages, 20 figures; AISTATS 2019
Cited by in corpus (14)
- How to Train Your Energy-Based Models
- Exponential Family Estimation via Adversarial Dynamics Embedding
- Training Deep Energy-Based Models with f-Divergence Minimization
- Noise Contrastive Meta-Learning for Conditional Density Estimation using Kernel Mean Embeddings
- Generalized Energy Based Models
- Energy-Inspired Models: Learning with Sampler-Induced Distributions
- Generative Modeling by Inclusive Neural Random Fields with Applications in Image Generation and Anomaly Detection
- Learning Discrete Energy-based Models via Auxiliary-variable Local Exploration
- Variational (Gradient) Estimate of the Score Function in Energy-based Latent Variable Models
- Scalable Personalised Item Ranking through Parametric Density Estimation
- Dual Training of Energy-Based Models with Overparametrized Shallow Neural Networks
- Convex-Concave Min-Max Stackelberg Games
- Kernel Deformed Exponential Families for Sparse Continuous Attention
- Denoising Score Matching with Random Fourier Features