A Mathematical Framework for Learning Probability Distributions
arXiv:2212.11481 · doi:10.4208/jml.221202
Abstract
The modeling of probability distributions, specifically generative modeling and density estimation, has become an immensely popular subject in recent years by virtue of its outstanding performance on sophisticated data such as images and texts. Nevertheless, a theoretical understanding of its success is still incomplete. One mystery is the paradox between memorization and generalization: In theory, the model is trained to be exactly the same as the empirical distribution of the finite samples, whereas in practice, the trained model can generate new samples or estimate the likelihood of unseen samples. Likewise, the overwhelming diversity of distribution learning models calls for a unified perspective on this subject. This paper provides a mathematical framework such that all the well-known models can be derived based on simple principles. To demonstrate its efficacy, we present a survey of our results on the approximation error, training error and generalization error of these models, which can all be established based on this framework. In particular, the aforementioned paradox is resolved by proving that these models enjoy implicit regularization during training, so that the generalization error at early-stopping avoids the curse of dimensionality. Furthermore, we provide some new results on landscape analysis and the mode collapse phenomenon.
fixed typos
References in corpus (25)
- Deep Learning in Neural Networks: An Overview
- Conditional Generative Adversarial Nets
- WaveNet: A Generative Model for Raw Audio
- Bootstrap your own latent: A new approach to self-supervised Learning
- PaLM: Scaling Language Modeling with Pathways
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- Score-Based Generative Modeling through Stochastic Differential Equations
- Zero-Shot Text-to-Image Generation
- Energy-based Generative Adversarial Network
- Generative Moment Matching Networks
- A Variational Approach to Enhanced Sampling and Free Energy Calculations
- Parallel WaveNet: Fast High-Fidelity Speech Synthesis
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- Nonparametric Density Estimation for High-Dimensional Data - Algorithms and Applications
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- A Priori Estimates of the Population Risk for Residual Networks
- Towards GAN Benchmarks Which Require Generalization
- Building Normalizing Flows with Stochastic Interpolants
- Understanding Overparameterization in Generative Adversarial Networks
- Alleviating Mode Collapse in GAN via Diversity Penalty Module
- Generalization and Memorization: The Bias Potential Model
- Unifying Generative Models with GFlowNets and Beyond
- The Slow Deterioration of the Generalization Error of the Random Feature Model
- Learning Optimal Flows for Non-Equilibrium Importance Sampling
- Optimal Transport Based Generative Autoencoders