Efficient Gradient-Based Inference through Transformations between Bayes Nets and Neural Nets
arXiv:1402.0480
Abstract
Hierarchical Bayesian networks and neural networks with stochastic hidden units are commonly perceived as two separate types of models. We show that either of these types of models can often be transformed into an instance of the other, by switching between centered and differentiable non-centered parameterizations of the latent variables. The choice of parameterization greatly influences the efficiency of gradient-based posterior inference; we show that they are often complementary to eachother, we clarify when each parameterization is preferred and show how inference can be made robust. In the non-centered form, a simple Monte Carlo estimator of the marginal likelihood can be used for learning the parameters. Theoretical results are supported by experiments.
References in corpus (4)
Cited by in corpus (13)
- Normalizing Flows for Probabilistic Modeling and Inference
- Gradient Estimation Using Stochastic Computation Graphs
- How Auto-Encoders Could Provide Credit Assignment in Deep Networks via Target Propagation
- Monte Carlo Gradient Estimation in Machine Learning
- Compositional generalization through meta sequence-to-sequence learning
- Reweighted Wake-Sleep
- GP-select: Accelerating EM using adaptive subspace preselection
- Big Learning with Bayesian Methods
- Note on Equivalence Between Recurrent Neural Network Time Series Models and Variational Bayesian Models
- MAP Propagation Algorithm: Faster Learning with a Team of Reinforcement Learning Agents
- Hybrid system identification using switching density networks
- Generative learning for deep networks
- Dropout Training for SVMs with Data Augmentation