Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
arXiv:1605.08104
Abstract
While great strides have been made in using deep learning algorithms to solve supervised learning tasks, the problem of unsupervised learning - leveraging unlabeled examples to learn about the structure of a domain - remains a difficult unsolved challenge. Here, we explore prediction of future frames in a video sequence as an unsupervised learning rule for learning about the structure of the visual world. We describe a predictive neural network ("PredNet") architecture that is inspired by the concept of "predictive coding" from the neuroscience literature. These networks learn to predict future frames in a video sequence, with each layer in the network making local predictions and only forwarding deviations from those predictions to subsequent network layers. We show that these networks are able to robustly learn to predict the movement of synthetic (rendered) objects, and that in doing so, the networks learn internal representations that are useful for decoding latent object parameters (e.g. pose) that support object recognition with fewer training views. We also show that these networks can scale to complex natural image streams (car-mounted camera videos), capturing key aspects of both egocentric movement and the movement of objects in the visual scene, and the representation learned in this setting is useful for estimating the steering angle. Altogether, these results suggest that prediction represents a powerful framework for unsupervised learning, allowing for implicit learning of object and scene structure.
Code and example video clips can be found here: https://coxlab.github.io/prednet/
References in corpus (24)
- Adam: A Method for Stochastic Optimization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- Deep Residual Learning for Image Recognition
- End to End Learning for Self-Driving Cars
- Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
- Going Deeper with Convolutions
- Deep Convolutional Inverse Graphics Network
- Semi-Supervised Learning with Ladder Networks
- Action-Conditional Video Prediction using Deep Networks in Atari Games
- Deep multi-scale video prediction beyond mean square error
- Video (language) modeling: a baseline for generative models of natural videos
- Unsupervised Learning for Physical Interaction through Video Prediction
- Spatio-temporal video autoencoder with differentiable memory
- Unsupervised Learning of Visual Representations using Videos
- Learning a Driving Simulator
- Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks
- Learning to See by Moving
- How Auto-Encoders Could Provide Credit Assignment in Deep Networks via Target Propagation
- Unsupervised Learning of Visual Structure using Predictive Generative Networks
- Learning Through Time in the Thalamocortical Loops
- Deep Predictive Coding Networks
- Understanding Visual Concepts with Continuation Learning
Cited by in corpus (138)
- Generating Videos with Scene Dynamics
- Convolutional Neural Networks as a Model of the Visual System: Past, Present, and Future
- Artificial neural networks for neuroscientists: A primer
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Unsupervised Learning for Physical Interaction through Video Prediction
- Stochastic Adversarial Video Prediction
- Task-Driven Convolutional Recurrent Models of the Visual System
- Contrastive Learning for Unpaired Image-to-Image Translation
- Video-to-Video Synthesis
- Space-Time Correspondence as a Contrastive Random Walk
- Train Sparsely, Generate Densely: Memory-efficient Unsupervised Training of High-resolution Temporal GAN
- Self-Supervised Visual Planning with Temporal Skip Connections
- Learning to Decompose and Disentangle Representations for Video Prediction
- Going in circles is the way forward: the role of recurrence in visual inference
- ODEVAE: Deep generative second order ODEs with Bayesian neural networks
- Convolutional Tensor-Train LSTM for Spatio-temporal Learning
- Spatiotemporal Contrastive Video Representation Learning
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
- Adaptive robot body learning and estimation through predictive coding
- Towards an integration of deep learning and neuroscience
- Stochastic Variational Video Prediction
- Future Frame Prediction for Anomaly Detection -- A New Baseline
- Small Sample Learning in Big Data Era
- Symbolic Pregression: Discovering Physical Laws from Distorted Video
- Less is More: Surgical Phase Recognition with Less Annotations through Self-Supervised Pre-training of CNN-LSTM Networks
- Deep Predictive Coding Network for Object Recognition
- A neural network walks into a lab: towards using deep nets as models for human behavior
- Crowding Reveals Fundamental Differences in Local vs. Global Processing in Humans and Machines
- Predictive-Corrective Networks for Action Detection
- End-to-end Learning of Driving Models from Large-scale Video Datasets
- A Taxonomy for Neural Memory Networks
- Cycle-Contrast for Self-Supervised Video Representation Learning
- Associative Memories via Predictive Coding
- Deep Predictive Learning: A Comprehensive Model of Three Visual Streams
- DDCNet: Deep Dilated Convolutional Neural Network for Dense Prediction
- Unsupervised Representation Learning by Sorting Sequences
- Colorful Image Colorization
- Deep Back-Projection Networks For Super-Resolution
- Predictive Coding: a Theoretical and Experimental Review
- Deep Learning Assisted Data Inspection for Radio Astronomy
- Predify: Augmenting deep neural networks with brain-inspired predictive coding dynamics
- Self-supervised Video Representation Learning by Uncovering Spatio-temporal Statistics
- Rules of the Road: Predicting Driving Behavior with a Convolutional Model of Semantic Interactions
- Machine Learning for the Detection and Identification of Internet of Things (IoT) Devices: A Survey
- Deep Visual Foresight for Planning Robot Motion
- Predictive Coding Can Do Exact Backpropagation on Convolutional and Recurrent Neural Networks
- Towards Learning to Detect and Predict Contact Events on Vision-based Tactile Sensors
- Learning Video Representations from Textual Web Supervision
- Inception-inspired LSTM for Next-frame Video Prediction
- Dual Motion GAN for Future-Flow Embedded Video Prediction
- Dynamics Transfer GAN: Generating Video by Transferring Arbitrary Temporal Dynamics from a Source Video to a Single Target Image
- Predictive Image Regression for Longitudinal Studies with Missing Data
- Deep Tracking on the Move: Learning to Track the World from a Moving Vehicle using Recurrent Neural Networks
- Prediction Under Uncertainty with Error-Encoding Networks
- Hybrid Learning of Optical Flow and Next Frame Prediction to Boost Optical Flow in the Wild
- Novel Video Prediction for Large-scale Scene using Optical Flow
- CortexNet: a Generic Network Family for Robust Visual Temporal Representations
- Unsupervised/Semi-supervised Deep Learning for Low-dose CT Enhancement
- Clockwork Variational Autoencoders
- Learning Bidirectional LSTM Networks for Synthesizing 3D Mesh Animation Sequences
- Motion Illusion-like Patterns Extracted from Photo and Art Images Using Predictive Deep Neural Networks
- A Deep Learning Based Attack for The Chaos-based Image Encryption
- Unsupervised Learning from Video with Deep Neural Embeddings
- Learning to Adapt by Minimizing Discrepancy
- Disentangling Propagation and Generation for Video Prediction
- Filter Grafting for Deep Neural Networks
- Self-Supervised Relative Depth Learning for Urban Scene Understanding
- VLUC: An Empirical Benchmark for Video-Like Urban Computing on Citywide Crowd and Traffic Prediction
- Few-shot Video-to-Video Synthesis
- Order Matters: Shuffling Sequence Generation for Video Prediction
- Distanced LSTM: Time-Distanced Gates in Long Short-Term Memory Models for Lung Cancer Detection
- Relational Action Forecasting
- CrevNet: Conditionally Reversible Video Prediction
- Stable and expressive recurrent vision models
- Greedy Hierarchical Variational Autoencoders for Large-Scale Video Prediction
- A Neurally-Inspired Hierarchical Prediction Network for Spatiotemporal Sequence Learning and Prediction
- EMPNet: Neural Localisation and Mapping Using Embedded Memory Points
- Adversarial Video Compression Guided by Soft Edge Detection
- Robust neural circuit reconstruction from serial electron microscopy with convolutional recurrent networks
- Learning Robot Activities from First-Person Human Videos Using Convolutional Future Regression
- Independent Innovation Analysis for Nonlinear Vector Autoregressive Process
- Sparse Coding on Stereo Video for Object Detection
- Multi-View Frame Reconstruction with Conditional GAN
- What Would You Do? Acting by Learning to Predict
- Generating Music Medleys via Playing Music Puzzle Games
- Pruning Filter in Filter
- Over-crowdedness Alert! Forecasting the Future Crowd Distribution
- Future Urban Scenes Generation Through Vehicles Synthesis
- Video Extrapolation with an Invertible Linear Embedding
- ContextVP: Fully Context-Aware Video Prediction
- A Temporally-Aware Interpolation Network for Video Frame Inpainting
- Sensorimotor Visual Perception on Embodied System Using Free Energy Principle
- Animating Landscape: Self-Supervised Learning of Decoupled Motion and Appearance for Single-Image Video Synthesis
- DDCNet-Multires: Effective Receptive Field Guided Multiresolution CNN for Dense Prediction
- Better Guider Predicts Future Better: Difference Guided Generative Adversarial Networks
- Predictive coding feedback results in perceived illusory contours in a recurrent neural network
- Binocular Rivalry Oriented Predictive Auto-Encoding Network for Blind Stereoscopic Image Quality Measurement
- On the difficulty of learning and predicting the long-term dynamics of bouncing objects
- Learning Temporal Dynamics from Cycles in Narrated Video
- Exploiting Spatio-Temporal Structure with Recurrent Winner-Take-All Networks
- An unsupervised long short-term memory neural network for event detection in cell videos
- Video Time: Properties, Encoders and Evaluation
- AFA-PredNet: The action modulation within predictive coding
- Better and Faster: Knowledge Transfer from Multiple Self-supervised Learning Tasks via Graph Distillation for Video Classification
- Neural Allocentric Intuitive Physics Prediction from Real Videos
- EvAn: Neuromorphic Event-based Anomaly Detection
- Learning Accurate Extended-Horizon Predictions of High Dimensional Trajectories
- Predictive coding, precision and natural gradients
- Artificial Perception Meets Psychophysics, Revealing a Fundamental Law of Illusory Motion
- CTNN: Corticothalamic-inspired neural network
- Reconstruction of Natural Visual Scenes from Neural Spikes with Deep Neural Networks
- Recurrent Flow-Guided Semantic Forecasting
- Flow Based Self-supervised Pixel Embedding for Image Segmentation
- Foresee: Attentive Future Projections of Chaotic Road Environments with Online Training
- Predicting the Future with Transformational States
- Forecasting Hands and Objects in Future Frames
- A neuro-inspired architecture for unsupervised continual learning based on online clustering and hierarchical predictive coding
- Unsupervised Learning Layers for Video Analysis
- A Framework for Multisensory Foresight for Embodied Agents
- Traversing Latent Space using Decision Ferns
- Adaptive Future Frame Prediction with Ensemble Network
- The many faces of deep learning
- Zero-shot Policy Learning with Spatial Temporal RewardDecomposition on Contingency-aware Observation
- Real-time Linear Operator Construction and State Estimation with the Kalman Filter
- Filter Grafting for Deep Neural Networks: Reason, Method, and Cultivation
- Spatio-Temporal Event Segmentation and Localization for Wildlife Extended Videos
- Exponential scaling of neural algorithms - a future beyond Moore's Law?
- Unsupervised Few-Shot Action Recognition via Action-Appearance Aligned Meta-Adaptation
- Predictive Coding Networks Meet Action Recognition
- Generation and Simulation of Yeast Microscopy Imagery with Deep Learning
- Learning Semantic-Aware Dynamics for Video Prediction
- Diversified Multiscale Graph Learning with Graph Self-Correction
- Variational Predictive Routing with Nested Subjective Timescales
- Recurrent networks improve neural response prediction and provide insights into underlying cortical circuits
- From Single to Multiple: Leveraging Multi-level Prediction Spaces for Video Forecasting
- Unsupervised predictive coding models may explain visual brain representation
- Hybrid Active Inference