On the ability of neural nets to express distributions
arXiv:1702.07028
Abstract
Deep neural nets have caused a revolution in many classification tasks. A related ongoing revolution -- also theoretically not understood -- concerns their ability to serve as generative models for complicated types of data such as images and texts. These models are trained using ideas like variational autoencoders and Generative Adversarial Networks. We take a first cut at explaining the expressivity of multilayer nets by giving a sufficient criterion for a function to be approximable by a neural network with hidden layers. A key ingredient is Barron's Theorem \cite{Barron1993}, which gives a Fourier criterion for approximability of a function by a neural network with 1 hidden layer. We show that a composition of functions which satisfy certain Fourier conditions ("Barron functions") can be approximated by a -layer neural network. For probability distributions, this translates into a criterion for a probability distribution to be approximable in Wasserstein distance -- a natural metric on probability distributions -- by a neural network applied to a fixed base distribution (e.g., multivariate gaussian). Building up recent lower bound work, we also give an example function that shows that composition of Barron functions is more expressive than Barron functions alone.
Accepted to COLT 2017
References in corpus (2)
Cited by in corpus (17)
- On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
- The jamming transition as a paradigm to understand the loss landscape of deep neural networks
- Efficient approximation of high-dimensional functions with neural networks
- On the capacity of deep generative networks for approximating distributions
- Porcupine Neural Networks: (Almost) All Local Optima are Global
- Greedy Layerwise Learning Can Scale to ImageNet
- Fourier Neural Networks: A Comparative Study
- An error analysis of generative adversarial networks for learning distributions
- On the Approximation Properties of Random ReLU Features
- How Powerful are Shallow Neural Networks with Bandlimited Random Weights?
- Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers
- A Deep Conditioning Treatment of Neural Networks
- Constructive Universal High-Dimensional Distribution Generation through Deep ReLU Networks
- Size-Noise Tradeoffs in Generative Networks
- Universal Approximation of Residual Flows in Maximum Mean Discrepancy
- Universal Joint Approximation of Manifolds and Densities by Simple Injective Flows
- A method for quantifying the generalization capabilities of generative models for solving Ising models