Early Visual Concept Learning with Unsupervised Deep Learning
arXiv:1606.05579
Abstract
Automated discovery of early visual concepts from raw image data is a major open challenge in AI research. Addressing this problem, we propose an unsupervised approach for learning disentangled representations of the underlying factors of variation. We draw inspiration from neuroscience, and show how this can be achieved in an unsupervised generative model by applying the same learning pressures as have been suggested to act in the ventral visual stream in the brain. By enforcing redundancy reduction, encouraging statistical independence, and exposure to data with transform continuities analogous to those to which human infants are exposed, we obtain a variational autoencoder (VAE) framework capable of learning disentangled factors. Our approach makes few assumptions and works well across a wide variety of datasets. Furthermore, our solution has useful emergent properties, such as zero-shot inference and an intuitive understanding of "objectness".
References in corpus (6)
- Deep Convolutional Inverse Graphics Network
- Weakly-supervised Disentangling with Recurrent Transformations for 3D View Synthesis
- Discovering Hidden Factors of Variation in Deep Networks
- Disentangling Factors of Variation via Generative Entangling
- High-Dimensional Probability Estimation with Deep Density Models
- Understanding Visual Concepts with Continuation Learning
Cited by in corpus (34)
- Deep Learning Approaches for Data Augmentation in Medical Imaging: A Review
- Towards Deep Symbolic Reinforcement Learning
- Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks
- Curiosity Driven Exploration of Learned Disentangled Goal Spaces
- Cognitive Psychology for Deep Neural Networks: A Shape Bias Case Study
- Near-Optimal Representation Learning for Hierarchical Reinforcement Learning
- Learning by Association - A versatile semi-supervised training method for neural networks
- Improving Generalization for Abstract Reasoning Tasks Using Disentangled Feature Representations
- Fixing a Broken ELBO
- Visual Dynamics: Stochastic Future Generation via Layered Cross Convolutional Networks
- FitVid: Overfitting in Pixel-Level Video Prediction
- Visual pathways from the perspective of cost functions and multi-task deep neural networks
- Probabilistic Model-Agnostic Meta-Learning
- Poincaré Wasserstein Autoencoder
- Disentangling Space and Time in Video with Hierarchical Variational Auto-encoders
- xGEMs: Generating Examplars to Explain Black-Box Models
- Learning by Abstraction: The Neural State Machine
- Latent Space Factorisation and Manipulation via Matrix Subspace Projection
- Learning to Predict Without Looking Ahead: World Models Without Forward Prediction
- Learning Style-Aware Symbolic Music Representations by Adversarial Autoencoders
- If deep learning is the answer, then what is the question?
- ToyArchitecture: Unsupervised Learning of Interpretable Models of the World
- Learning Potentials of Quantum Systems using Deep Neural Networks
- Semantics Preserving Adversarial Learning
- Variational Collaborative Learning for User Probabilistic Representation
- Semi-Supervised Learning with the Deep Rendering Mixture Model
- Efficient Convolutional Auto-Encoding via Random Convexification and Frequency-Domain Minimization
- Generalization to Novel Objects using Prior Relational Knowledge
- Proximity Variational Inference
- Unsupervised Learning Layers for Video Analysis
- Autonomous Goal Exploration using Learned Goal Spaces for Visuomotor Skill Acquisition in Robots
- Robust Disentanglement of a Few Factors at a Time
- Learning Good Representation via Continuous Attention
- Quantifying the Effects of Enforcing Disentanglement on Variational Autoencoders