Deep Generative Stochastic Networks Trainable by Backprop
arXiv:1306.1091
Abstract
We introduce a novel training principle for probabilistic models that is an alternative to maximum likelihood. The proposed Generative Stochastic Networks (GSN) framework is based on learning the transition operator of a Markov chain whose stationary distribution estimates the data distribution. The transition distribution of the Markov chain is conditional on the previous state, generally involving a small move, so this conditional distribution has fewer dominant modes, being unimodal in the limit of small moves. Thus, it is easier to learn because it is easier to approximate its partition function, more like learning to perform supervised function approximation, with gradients that can be obtained by backprop. We provide theorems that generalize recent work on the probabilistic interpretation of denoising autoencoders and obtain along the way an interesting justification for dependency networks and generalized pseudolikelihood, along with a definition of an appropriate joint distribution and sampling mechanism even when the conditionals are not consistent. GSNs can be used with missing inputs and can be used to sample subsets of variables given the rest. We validate these theoretical results with experiments on two image datasets using an architecture that mimics the Deep Boltzmann Machine Gibbs sampler but allows training to proceed with simple backprop, without the need for layerwise pretraining.
arXiv admin note: text overlap with arXiv:1305.0445, Also published in ICML'2014
References in corpus (9)
- Improving neural networks by preventing co-adaptation of feature detectors
- Stochastic Backpropagation and Approximate Inference in Deep Generative Models
- Sum-Product Networks: A New Deep Architecture
- Better Mixing via Deep Representations
- Generalized Denoising Auto-Encoders as Generative Models
- Estimating or Propagating Gradients Through Stochastic Neurons
- Learning Feature Hierarchies with Centered Deep Boltzmann Machines
- A Generative Process for Sampling Contractive Auto-Encoders
- Fast Gradient-Based Inference with Continuous Latent Variable Models in Auxiliary Form
Cited by in corpus (80)
- An Introduction to Variational Autoencoders
- Deep Unsupervised Learning using Nonequilibrium Thermodynamics
- Tutorial on Variational Autoencoders
- NIPS 2016 Tutorial: Generative Adversarial Networks
- Adversarially Learned Inference
- A note on the evaluation of generative models
- Generative Moment Matching Networks
- Learning with Pseudo-Ensembles
- MADE: Masked Autoencoder for Distribution Estimation
- Unsupervised and Semi-supervised Learning with Categorical Generative Adversarial Networks
- Recent Advances in Convolutional Neural Networks
- Adversarial Autoencoders
- Towards Biologically Plausible Deep Learning
- Unsupervised Visual Representation Learning by Context Prediction
- Opportunities and challenges for quantum-assisted machine learning in near-term quantum computers
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- Image De-raining Using a Conditional Generative Adversarial Network
- Generalized Denoising Auto-Encoders as Generative Models
- Discovering Hidden Factors of Variation in Deep Networks
- Generative Visual Manipulation on the Natural Image Manifold
- Deep Learning of Representations: Looking Forward
- Denoising Diffusion Implicit Models
- Attribute2Image: Conditional Image Generation from Visual Attributes
- Unsupervised Learning of View-invariant Action Representations
- Super-Resolution with Deep Convolutional Sufficient Statistics
- Monte Carlo Gradient Estimation in Machine Learning
- DESIRE: Distant Future Prediction in Dynamic Scenes with Interacting Agents
- A-NICE-MC: Adversarial Training for MCMC
- Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends
- Variational Auto-encoded Deep Gaussian Processes
- Implicit Maximum Likelihood Estimation
- Learning Implicit Fields for Generative Shape Modeling
- Why are deep nets reversible: A simple theory, with implications for training
- Efficient Gradient-Based Inference through Transformations between Bayes Nets and Neural Nets
- Multi-view Generative Adversarial Networks
- Automatic recognition of element classes and boundaries in the birdsong with variable sequences
- Creativity: Generating Diverse Questions using Variational Autoencoders
- A Study of the Generalizability of Self-Supervised Representations
- Generalizing Hamiltonian Monte Carlo with Neural Networks
- Transport Analysis of Infinitely Deep Neural Network
- Learning with hidden variables
- Max-margin Deep Generative Models
- Learning to Generate Chairs, Tables and Cars with Convolutional Networks
- Tensorial Mixture Models
- Bounding the Test Log-Likelihood of Generative Models
- Auto-Differentiating Linear Algebra
- Learning to Generate with Memory
- Top-Down Learning for Structured Labeling with Convolutional Pseudoprior
- Inference in Deep Networks in High Dimensions
- F-Divergences and Cost Function Locality in Generative Modelling with Quantum Circuits
- AUC-maximized Deep Convolutional Neural Fields for Sequence Labeling
- Natural Language Generation with Neural Variational Models
- Predicting distributions with Linearizing Belief Networks
- Multimodal Transitions for Generative Stochastic Networks
- Image Generation and Editing with Variational Info Generative AdversarialNetworks
- Learning Deep Generative Models with Doubly Stochastic MCMC
- Locally Masked Convolution for Autoregressive Models
- Variational Generative Stochastic Networks with Collaborative Shaping
- Visual Language Modeling on CNN Image Representations
- Cooperative Training of Fast Thinking Initializer and Slow Thinking Solver for Conditional Learning
- The Labeling Distribution Matrix (LDM): A Tool for Estimating Machine Learning Algorithm Capacity
- Improved Autoregressive Modeling with Distribution Smoothing
- Tagger: Deep Unsupervised Perceptual Grouping
- Automatic Validation of Textual Attribute Values in E-commerce Catalog by Learning with Limited Labeled Data
- Generative Mixture of Networks
- Understanding Dropout: Training Multi-Layer Perceptrons with Auxiliary Independent Stochastic Neurons
- Symmetries and control in generative neural nets
- Population-based Gradient Descent Weight Learning for Graph Coloring Problems
- Max-Margin Deep Generative Models for (Semi-)Supervised Learning
- AdvNF: Reducing Mode Collapse in Conditional Normalising Flows using Adversarial Learning
- Harnessing Optoelectronic Noises in a Photonic Generative Network
- Deep Secure Encoding: An Application to Face Recognition
- Convex Smoothed Autoencoder-Optimal Transport model
- On educating machines
- Revealing the Distributional Vulnerability of Discriminators by Implicit Generators
- Towards GANs' Approximation Ability
- Coresets for Dependency Networks
- Adiabatic Persistent Contrastive Divergence Learning
- On the Generative Utility of Cyclic Conditionals
- Image Disguise based on Generative Model