Density Estimation in Infinite Dimensional Exponential Families
arXiv:1312.3516
Abstract
In this paper, we consider an infinite dimensional exponential family, of probability densities, which are parametrized by functions in a reproducing kernel Hilbert space, and show it to be quite rich in the sense that a broad class of densities on can be approximated arbitrarily well in Kullback-Leibler (KL) divergence by elements in . The main goal of the paper is to estimate an unknown density, through an element in . Standard techniques like maximum likelihood estimation (MLE) or pseudo MLE (based on the method of sieves), which are based on minimizing the KL divergence between and , do not yield practically useful estimators because of their inability to efficiently handle the log-partition function. Instead, we propose an estimator, based on minimizing the \emph{Fisher divergence}, between and , which involves solving a simple finite-dimensional linear system. When , we show that the proposed estimator is consistent, and provide a convergence rate of in Fisher divergence under the smoothness assumption that for some , where is a certain Hilbert-Schmidt operator on and denotes the image of . We also investigate the misspecified case of and show that as , and provide a rate for this convergence under a similar smoothness condition as above. Through numerical simulations we demonstrate that the proposed estimator outperforms the non-parametric kernel density estimator, and that the advantage with the proposed estimator grows as increases.
58 pages, 8 figures; Fixed some errors and typos
Cited by in corpus (39)
- Kernel Mean Embedding of Distributions: A Review and Beyond
- Learning Decentralized Controllers for Robot Swarms with Graph Neural Networks
- Learning the PE Header, Malware Detection with Minimal Domain Knowledge
- Beating the Perils of Non-Convexity: Guaranteed Training of Neural Networks using Tensor Methods
- Variational Autoencoders and Nonlinear ICA: A Unifying Framework
- A Kernel Test of Goodness of Fit
- Sliced Score Matching: A Scalable Approach to Density and Score Estimation
- Optimal Rates for Random Fourier Features
- Kernel Mean Shrinkage Estimators
- Learning Theory for Distribution Regression
- Gradient-free Hamiltonian Monte Carlo with Efficient Kernel Exponential Families
- Spectral Learning on Matrices and Tensors
- Minimum Stein Discrepancy Estimators
- Exponential Family Estimation via Adversarial Dynamics Embedding
- Generalization Properties of Optimal Transport GANs with Latent Distribution Learning
- A Kernelized Stein Discrepancy for Goodness-of-fit Tests and Model Evaluation
- Training Input-Output Recurrent Neural Networks through Spectral Methods
- Learning deep kernels for exponential family densities
- Generalization and Memorization: The Bias Potential Model
- Kernelized Wasserstein Natural Gradient
- Generalized Energy Based Models
- PSD Representations for Effective Probability Models
- Efficient and principled score estimation with Nyström kernel exponential families
- Efficient Learning of Generative Models via Finite-Difference Score Matching
- Identifying through Flows for Recovering Latent Representations
- Scalable Gaussian Process Inference with Finite-data Mean and Variance Guarantees
- Learning high-dimensional graphical models with regularized quadratic scoring
- Kernel Conditional Density Operators
- Estimating linear response statistics using orthogonal polynomials: An RKHS formulation
- The Cramér-Rao inequality on singular statistical models I
- Nonparametric Score Estimators
- Active Slices for Sliced Stein Discrepancy
- A Wasserstein Minimum Velocity Approach to Learning Unnormalized Models
- Kernel Deformed Exponential Families for Sparse Continuous Attention
- Scalable Personalised Item Ranking through Parametric Density Estimation
- MLE convergence speed to information projection of exponential family: Criterion for model dimension and sample size -- complete proof version--
- Weighting-Based Treatment Effect Estimation via Distribution Learning
- Unsupervised and Supervised Structure Learning for Protein Contact Prediction
- A Statistical Taylor Theorem and Extrapolation of Truncated Densities