A Deep and Tractable Density Estimator
arXiv:1310.1757
Abstract
The Neural Autoregressive Distribution Estimator (NADE) and its real-valued version RNADE are competitive density models of multidimensional data across a variety of domains. These models use a fixed, arbitrary ordering of the data dimensions. One can easily condition on variables at the beginning of the ordering, and marginalize out variables at the end of the ordering, however other inference tasks require approximate inference. In this work we introduce an efficient procedure to simultaneously train a NADE model for each possible ordering of the variables, by sharing parameters across all these models. We can thus use the most convenient model for each inference task at hand, and ensembles of such models with different orderings are immediately available. Moreover, unlike the original NADE, our training procedure scales to deep models. Empirically, ensembles of Deep NADE models obtain state of the art density estimation performance.
9 pages, 4 tables, 1 algorithm, 5 figures. To appear ICML 2014, JMLR W&CP volume 32
References in corpus (2)
Cited by in corpus (51)
- Pixel Recurrent Neural Networks
- Variational Inference with Normalizing Flows
- DRAW: A Recurrent Neural Network For Image Generation
- Markov Chain Monte Carlo and Variational Inference: Bridging the Gap
- Normalizing Flows for Probabilistic Modeling and Inference
- MADE: Masked Autoencoder for Distribution Estimation
- Deep Unsupervised Cardinality Estimation
- Masked Autoregressive Flow for Density Estimation
- A Neural Autoregressive Approach to Collaborative Filtering
- Deep Learning-Based Video Coding: A Review and A Case Study
- How Auto-Encoders Could Provide Credit Assignment in Deep Networks via Target Propagation
- Deep AutoRegressive Networks
- Generative Image Modeling Using Spatial LSTMs
- Nonparametric Density Estimation for High-Dimensional Data - Algorithms and Applications
- Counterpoint by Convolution
- A Deep and Autoregressive Approach for Topic Modeling of Multimodal Data
- Variational Auto-encoded Deep Gaussian Processes
- The Bach Doodle: Approachable music composition with machine learning at scale
- SurVAE Flows: Surjections to Bridge the Gap between VAEs and Flows
- Conditional Density Estimation with Bayesian Normalising Flows
- IDF++: Analyzing and Improving Integer Discrete Flows for Lossless Compression
- Reweighted Wake-Sleep
- Perfect density models cannot guarantee anomaly detection
- Iterative Neural Autoregressive Distribution Estimator (NADE-k)
- Deep Directed Generative Autoencoders
- Transformation Autoregressive Networks
- Bidirectional Recurrent Neural Networks as Generative Models - Reconstructing Gaps in Time Series
- Neural Likelihoods via Cumulative Distribution Functions
- TraDE: Transformers for Density Estimation
- Efficient Bayesian Inference for a Gaussian Process Density Model
- Diagnostics for Conditional Density Models and Bayesian Inference Algorithms
- LogitBoost autoregressive networks
- Document Neural Autoregressive Distribution Estimation
- Distilling Model Knowledge
- Learning Deep Generative Models with Doubly Stochastic MCMC
- Learning about an exponential amount of conditional distributions
- Locally Masked Convolution for Autoregressive Models
- Generating Music with a Self-Correcting Non-Chronological Autoregressive Model
- DeepSPACE: Approximate Geospatial Query Processing with Deep Learning
- Filtering Variational Objectives
- Representation Learning for Remote Sensing: An Unsupervised Sensor Fusion Approach
- Mixtures of Sparse Autoregressive Networks
- Melody Harmonization Using Orderless NADE, Chord Balancing, and Blocked Gibbs Sampling
- Soft-Deep Boltzmann Machines
- The DEformer: An Order-Agnostic Distribution Estimating Transformer
- Autoregressive Diffusion Models
- Arbitrary Conditional Distributions with Energy
- Recurrent Estimation of Distributions
- Predicting Classification Accuracy When Adding New Unobserved Classes
- Towards Representation Learning with Tractable Probabilistic Models
- An Improved Training Procedure for Neural Autoregressive Data Completion