Unsupervised Learning of Visual Structure using Predictive Generative Networks
arXiv:1511.06380
Abstract
The ability to predict future states of the environment is a central pillar of intelligence. At its core, effective prediction requires an internal model of the world and an understanding of the rules by which the world changes. Here, we explore the internal models developed by deep neural networks trained using a loss based on predicting future frames in synthetic video sequences, using a CNN-LSTM-deCNN framework. We first show that this architecture can achieve excellent performance in visual sequence prediction tasks, including state-of-the-art performance in a standard 'bouncing balls' dataset (Sutskever et al., 2009). Using a weighted mean-squared error and adversarial loss (Goodfellow et al., 2014), the same architecture successfully extrapolates out-of-the-plane rotations of computer-generated faces. Furthermore, despite being trained end-to-end to predict only pixel-level information, our Predictive Generative Networks learn a representation of the latent structure of the underlying three-dimensional objects themselves. Importantly, we find that this representation is naturally tolerant to object transformations, and generalizes well to new tasks, such as classification of static images. Similar models trained solely with a reconstruction loss fail to generalize as effectively. We argue that prediction can serve as a powerful unsupervised loss for learning rich internal representations of high-level object features.
under review as conference paper at ICLR 2016
References in corpus (6)
- Sequence to Sequence Learning with Neural Networks
- Conditional Generative Adversarial Nets
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- Deep Convolutional Inverse Graphics Network
- Deep Predictive Coding Networks
- Predictive Encoding of Contextual Relationships for Perceptual Inference, Interpolation and Prediction
Cited by in corpus (23)
- NIPS 2016 Tutorial: Generative Adversarial Networks
- Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
- A Review on Deep Learning Techniques for Video Prediction
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- ODEVAE: Deep generative second order ODEs with Bayesian neural networks
- SpaceNet MVOI: a Multi-View Overhead Imagery Dataset
- Towards Adversarial Retinal Image Synthesis
- Towards an integration of deep learning and neuroscience
- Subsurface structure analysis using computational interpretation and learning: A visual signal processing perspective
- Void Filling of Digital Elevation Models with Deep Generative Models
- Synthesizing realistic neural population activity patterns using Generative Adversarial Networks
- Generative Adversarial Networks: A Survey Towards Private and Secure Applications
- Global-Local Face Upsampling Network
- A neural network trained to predict future video frames mimics critical properties of biological neuronal responses and perception
- Inception-inspired LSTM for Next-frame Video Prediction
- Unsupervised Learning from Continuous Video in a Scalable Predictive Recurrent Network
- End-to-End Adaptive Monte Carlo Denoising and Super-Resolution
- On the difficulty of learning and predicting the long-term dynamics of bouncing objects
- Multi Resolution LSTM For Long Term Prediction In Neural Activity Video
- Hierarchical View Predictor: Unsupervised 3D Global Feature Learning through Hierarchical Prediction among Unordered Views
- Enhancing audio quality for expressive Neural Text-to-Speech
- Point-to-Point Video Generation
- Deep Variational Luenberger-type Observer for Stochastic Video Prediction