MADE: Masked Autoencoder for Distribution Estimation
arXiv:1502.03509
Abstract
There has been a lot of recent interest in designing neural network models to estimate a distribution from a set of examples. We introduce a simple modification for autoencoder neural networks that yields powerful generative models. Our method masks the autoencoder's parameters to respect autoregressive constraints: each input is reconstructed only from previous inputs in a given ordering. Constrained this way, the autoencoder outputs can be interpreted as a set of conditional probabilities, and their product, the full joint probability. We can also train a single network that can decompose the joint probability in multiple different orderings. Our simple framework can be applied to multiple architectures, including deep ones. Vectorized implementations, such as on GPUs, are simple and fast. Experiments demonstrate that this approach is competitive with state-of-the-art tractable distribution estimators. At test time, the method is significantly faster and scales better than other autoregressive estimators.
9 pages and 1 page of supplementary material. Updated to match published version
References in corpus (8)
- Auto-Encoding Variational Bayes
- ADADELTA: An Adaptive Learning Rate Method
- Theano: new features and speed improvements
- Sum-Product Networks: A New Deep Architecture
- Deep Generative Stochastic Networks Trainable by Backprop
- Deep AutoRegressive Networks
- A Deep and Tractable Density Estimator
- Learning Representations by Maximizing Compression
Cited by in corpus (201)
- An Introduction to Variational Autoencoders
- XLNet: Generalized Autoregressive Pretraining for Language Understanding
- Pixel Recurrent Neural Networks
- Normalizing Flows: An Introduction and Review of Current Methods
- The frontier of simulation-based inference
- Density estimation using Real NVP
- Normalizing Flows for Probabilistic Modeling and Inference
- Professor Forcing: A New Algorithm for Training Recurrent Networks
- Anomaly Detection with Density Estimation
- GraphAF: a Flow-based Autoregressive Model for Molecular Graph Generation
- Deep Unsupervised Cardinality Estimation
- Masked Autoregressive Flow for Density Estimation
- Improving Variational Inference with Inverse Autoregressive Flow
- Gravitational-wave parameter estimation with autoregressive neural network flows
- Mining gold from implicit models to improve likelihood-free inference
- Differentiable Quantum Architecture Search
- Invertible Residual Networks
- Neural Network Renormalization Group
- FFJORD: Free-form Continuous Dynamics for Scalable Reversible Generative Models
- The Dark Machines Anomaly Score Challenge: Benchmark Data and Model Independent Event Classification for the Large Hadron Collider
- From Variational to Deterministic Autoencoders
- Roundtrip: A Deep Generative Neural Density Estimator
- Nonparametric Density Estimation for High-Dimensional Data - Algorithms and Applications
- How to Train Your Energy-Based Models
- A topological insight into restricted Boltzmann machines
- Mining for Dark Matter Substructure: Inferring subhalo population properties from strong lenses with machine learning
- Bayesian Layers: A Module for Neural Network Uncertainty
- Equivariant Flows: Exact Likelihood Generative Learning for Symmetric Densities
- Sequential Neural Likelihood: Fast Likelihood-free Inference with Autoregressive Flows
- PolyGen: An Autoregressive Generative Model of 3D Meshes
- Non-Gaussian estimates of tensions in cosmological parameters
- Variational Autoencoder with Arbitrary Conditioning
- Rare and Different: Anomaly Scores from a combination of likelihood and out-of-distribution models to detect new physics at the LHC
- Inference of the optical depth to reionization from low multipole temperature and polarisation Planck data
- Low-Quality Training Data Only? A Robust Framework for Detecting Encrypted Malicious Network Traffic
- ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis
- Inference Networks for Sequential Monte Carlo in Graphical Models
- Stacked Generative Adversarial Networks
- Targeted free energy perturbation revisited: Accurate free energies from mapped reference potentials
- Monge-Ampère Flow for Generative Modeling
- Dynamical mass inference of galaxy clusters with neural flows
- Fixing a Broken ELBO
- Copula Flows for Synthetic Data Generation
- Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions
- Tensor networks for unsupervised machine learning
- The sum of the masses of the Milky Way and M31: a likelihood-free inference approach
- Policy Gradient based Quantum Approximate Optimization Algorithm
- A Tutorial on Deep Latent Variable Models of Natural Language
- Simulation-based inference of dynamical galaxy cluster masses with 3D convolutional neural networks
- Pixel Deconvolutional Networks
- CaloFlow II: Even Faster and Still Accurate Generation of Calorimeter Showers with Normalizing Flows
- Cross-Camera Feature Prediction for Intra-Camera Supervised Person Re-identification across Distant Scenes
- Likelihood-free MCMC with Amortized Approximate Ratio Estimators
- Neural Spline Flows
- Neural Manifold Ordinary Differential Equations
- NeuroCard: One Cardinality Estimator for All Tables
- Block Neural Autoregressive Flow
- Calculating Renyi Entropies with Neural Autoregressive Quantum States
- SARM: Sparse Autoregressive Model for Scalable Generation of Sparse Images in Particle Physics
- Ephemeral Learning -- Augmenting Triggers with Online-Trained Normalizing Flows
- Relaxing Bijectivity Constraints with Continuously Indexed Normalising Flows
- Effective LHC measurements with matrix elements and machine learning
- Benchmarking deep inverse models over time, and the neural-adjoint method
- Multi-Attribute Selectivity Estimation Using Deep Learning
- Multimap targeted free energy estimation
- Neural Autoregressive Distribution Estimation
- Image-to-Image Translation: Methods and Applications
- Learnable Explicit Density for Continuous Latent Space and Variational Inference
- Transformation Autoregressive Networks
- Cardinality Estimation in DBMS: A Comprehensive Benchmark Evaluation
- Variational Inference with NoFAS: Normalizing Flow with Adaptive Surrogate for Computationally Expensive Models
- Why Not to Use Zero Imputation? Correcting Sparsity Bias in Training Neural Networks
- of two-dimensional electron gas: a neural canonical transformation study
- Gradient-Based Neural DAG Learning
- Bayesian Hypernetworks
- Degeneration in VAE: in the Light of Fisher Information Loss
- Scalable inference with Autoregressive Neural Ratio Estimation
- Constructing Deep Neural Networks by Bayesian Network Structure Learning
- PixelVAE++: Improved PixelVAE with Discrete Prior
- MAE: Mutual Posterior-Divergence Regularization for Variational AutoEncoders
- TraDE: Transformers for Density Estimation
- Neural Likelihoods via Cumulative Distribution Functions
- Deep Generative Quantile-Copula Models for Probabilistic Forecasting
- Neural canonical transformations for vibrational spectra of molecules
- The Causal-Neural Connection: Expressiveness, Learnability, and Inference
- MultiVerse: Causal Reasoning using Importance Sampling in Probabilistic Programming
- Graphical Normalizing Flows
- Low-rank Characteristic Tensor Density Estimation Part I: Foundations
- Training Deep Energy-Based Models with f-Divergence Minimization
- Convolutional Normalizing Flows
- Counterfactual Data Augmentation using Locally Factored Dynamics
- Interventional Sum-Product Networks: Causal Inference with Tractable Probabilistic Models
- Neural Empirical Bayes: Source Distribution Estimation and its Applications to Simulation-Based Inference
- Unconstrained Monotonic Neural Networks
- Generalizing to new geometries with Geometry-Aware Autoregressive Models (GAAMs) for fast calorimeter simulation
- Visualizing and Understanding Sum-Product Networks
- The Neural Moving Average Model for Scalable Variational Inference of State Space Models
- Comparison of Affine and Rational Quadratic Spline Coupling and Autoregressive Flows through Robust Statistical Tests
- Low-rank Characteristic Tensor Density Estimation Part II: Compression and Latent Density Estimation
- Noise-Robust End-to-End Quantum Control using Deep Autoregressive Policy Networks
- Invertible Generative Modeling using Linear Rational Splines
- BayesCard: Revitilizing Bayesian Frameworks for Cardinality Estimation
- LogitBoost autoregressive networks
- Convex Potential Flows: Universal Probability Distributions with Optimal Transport and Convex Optimization
- Faithful Inversion of Generative Models for Effective Amortized Inference
- Neural Recursive Belief States in Multi-Agent Reinforcement Learning
- Gaussianization Flows
- Estimating Photometric Redshifts for Galaxies from the DESI Legacy Imaging Surveys with Bayesian Neural Networks Trained by DESI EDR
- Locally Masked Convolution for Autoregressive Models
- Measurement of the and cross-sections in collisions at TeV with the ATLAS detector
- EMaQ: Expected-Max Q-Learning Operator for Simple Yet Effective Offline and Online RL
- Predictive Sampling with Forecasting Autoregressive Models
- Trial by FIRE: Probing the dark matter density profile of dwarf galaxies with GraphNPE
- Fast and Flexible Temporal Point Processes with Triangular Maps
- Dual Supervised Learning for Natural Language Understanding and Generation
- What does a cosmological experiment really measure? Covariant posterior decomposition with normalizing flows
- Adaptive Monte Carlo augmented with normalizing flows
- Modeling the Gaia Color-Magnitude Diagram with Bayesian Neural Flows to Constrain Distance Estimates
- Anytime Sampling for Autoregressive Models via Ordered Autoencoding
- Language Modeling with Reduced Densities
- Solving Quantum Statistical Mechanics with Variational Autoregressive Networks and Quantum Circuits
- Sequential Neural Methods for Likelihood-free Inference
- A Unified Transferable Model for ML-Enhanced DBMS
- HNPE: Leveraging Global Parameters for Neural Posterior Estimation
- Towards Unsupervised Language Understanding and Generation by Joint Dual Learning
- MaCow: Masked Convolutional Generative Flow
- Data-driven deep density estimation
- Trust-Region Variational Inference with Gaussian Mixture Models
- Probabilistic Software Modeling: A Data-driven Paradigm for Software Analysis
- DeepSPACE: Approximate Geospatial Query Processing with Deep Learning
- Improved Autoregressive Modeling with Distribution Smoothing
- Generalized Energy Based Models
- Sparse Flows: Pruning Continuous-depth Models
- Annealed Flow Transport Monte Carlo
- Automating Inference of Binary Microlensing Events with Neural Density Estimation
- Backpropagation for Implicit Spectral Densities
- A novel stellar spectrum denoising method based on deep Bayesian modeling
- The Convolution Exponential and Generalized Sylvester Flows
- Robust normalizing flows using Bernstein-type polynomials
- Featurized Density Ratio Estimation
- Structured Sparsity Inducing Adaptive Optimizers for Deep Learning
- Autoregressive Score Matching
- Hyperparameter optimization with REINFORCE and Transformers
- A lower bound for the ELBO of the Bernoulli Variational Autoencoder
- Fair Normalizing Flows
- Woodbury Transformations for Deep Generative Flows
- Closing the Dequantization Gap: PixelCNN as a Single-Layer Flow
- Measure Transport with Kernel Stein Discrepancy
- General Probabilistic Surface Optimization and Log Density Estimation
- A Forest from the Trees: Generation through Neighborhoods
- Black-Box Autoregressive Density Estimation for State-Space Models
- Auto-decoding Graphs
- Set Distribution Networks: a Generative Model for Sets of Images
- Mixtures of Sparse Autoregressive Networks
- Uncertainty-aware Cardinality Estimation by Neural Network Gaussian Process
- Black-Box Inference for Non-Linear Latent Force Models
- Learning Markov Random Fields for Combinatorial Structures via Sampling through Lovász Local Lemma
- Arbitrary Conditional Distributions with Energy
- Copula-like Variational Inference
- Masking schemes for universal marginalisers
- Causal Autoregressive Flows
- Re-examination of the Role of Latent Variables in Sequence Modeling
- Training Invertible Linear Layers through Rank-One Perturbations
- Learning summary features of time series for likelihood free inference
- TzK Flow - Conditional Generative Model
- Deep PDF: Probabilistic Surface Optimization and Density Estimation
- Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word Order
- Universal Approximation of Residual Flows in Maximum Mean Discrepancy
- Towards Recurrent Autoregressive Flow Models
- Dual Inference for Improving Language Understanding and Generation
- Density Deconvolution with Normalizing Flows
- Reliable Categorical Variational Inference with Mixture of Discrete Normalizing Flows
- Universal Marginaliser for Deep Amortised Inference for Probabilistic Programs
- Gradient Boosted Normalizing Flows
- Semi-Implicit Stochastic Recurrent Neural Networks
- Hierarchical Reinforcement Learning with Optimal Level Synchronization Based on Flow-Based Deep Generative Model
- Efficient sampling generation from explicit densities via Normalizing Flows
- Variational Autoencoded Regression: High Dimensional Regression of Visual Data on Complex Manifold
- Lossless Compression with Latent Variable Models
- Robust Audio Anomaly Detection
- Coping With Simulators That Don't Always Return
- Modeling Diagnostic Label Correlation for Automatic ICD Coding
- A Triangular Network For Density Estimation
- TzK: Flow-Based Conditional Generative Model
- Automated Dependence Plots
- Diffusion Normalizing Flow
- Sinusoidal Flow: A Fast Invertible Autoregressive Flow
- Invertible Attention
- Imitation with Neural Density Models
- Gradient-based Causal Structure Learning with Normalizing Flow
- Discrete Tree Flows via Tree-Structured Permutations
- Self-Reflective Variational Autoencoder
- Overcoming barriers to scalability in variational quantum Monte Carlo
- Synthetic Data Generation for Economists
- GACEM: Generalized Autoregressive Cross Entropy Method for Multi-Modal Black Box Constraint Satisfaction
- Molecular Attributes Transfer from Non-Parallel Data
- High Mutual Information in Representation Learning with Symmetric Variational Inference
- Discriminative, Generative and Self-Supervised Approaches for Target-Agnostic Learning
- Neural Approximation of an Auto-Regressive Process through Confidence Guided Sampling
- Probabilistic Autoencoder using Fisher Information
- Weighting-Based Treatment Effect Estimation via Distribution Learning