Neural Simpletrons - Minimalistic Directed Generative Networks for Learning with Few Labels
arXiv:1506.08448 · doi:10.1162/neco_a_01100
Abstract
Classifiers for the semi-supervised setting often combine strong supervised models with additional learning objectives to make use of unlabeled data. This results in powerful though very complex models that are hard to train and that demand additional labels for optimal parameter tuning, which are often not given when labeled data is very sparse. We here study a minimalistic multi-layer generative neural network for semi-supervised learning in a form and setting as similar to standard discriminative networks as possible. Based on normalized Poisson mixtures, we derive compact and local learning and neural activation rules. Learning and inference in the network can be scaled using standard deep learning tools for parallelized GPU implementation. With the single objective of likelihood optimization, both labeled and unlabeled data are naturally incorporated into learning. Empirical evaluations on standard benchmarks show, that for datasets with few labels the derived minimalistic network improves on all classical deep learning approaches and is competitive with their recent variants without the need of additional labels for parameter tuning. Furthermore, we find that the studied network is the best performing monolithic (`non-hybrid') system for few labels, and that it can be applied in the limit of very few labels, where no other system has been reported to operate so far.
References in corpus (10)
- Adam: A Method for Stochastic Optimization
- Deep Learning in Neural Networks: An Overview
- Bayesian Convolutional Neural Networks with Bernoulli Approximate Variational Inference
- Semi-Supervised Learning with Ladder Networks
- Deep Bayesian Active Learning with Image Data
- Distributional Smoothing with Virtual Adversarial Training
- Deep Networks with Stochastic Depth
- GP-select: Accelerating EM using adaptive subspace preselection
- Truncated Variational Expectation Maximization
- Truncated Variational EM for Semi-Supervised Neural Simpletrons
Cited by in corpus (6)
- An Overview of Deep Semi-Supervised Learning
- Realistic Evaluation of Deep Semi-Supervised Learning Algorithms
- Regularization by architecture: A deep prior approach for inverse problems
- -means as a variational EM approximation of Gaussian mixture models
- Large Scale Clustering with Variational EM for Gaussian Mixture Models
- Evolutionary Variational Optimization of Generative Models