Stochastic Variational Video Prediction
arXiv:1710.11252
Abstract
Predicting the future in real-world settings, particularly from raw sensory observations such as images, is exceptionally challenging. Real-world events can be stochastic and unpredictable, and the high dimensionality and complexity of natural images requires the predictive model to build an intricate understanding of the natural world. Many existing methods tackle this problem by making simplifying assumptions about the environment. One common assumption is that the outcome is deterministic and there is only one plausible future. This can lead to low-quality predictions in real-world settings with stochastic dynamics. In this paper, we develop a stochastic variational video prediction (SV2P) method that predicts a different possible future for each sample of its latent variables. To the best of our knowledge, our model is the first to provide effective stochastic multi-frame prediction for real-world video. We demonstrate the capability of the proposed method in predicting detailed future frames of videos on multiple real-world datasets, both action-free and action-conditioned. We find that our proposed method produces substantially improved video predictions when compared to the same model without stochasticity, and to other stochastic video prediction methods. Our SV2P implementation will be open sourced upon publication.
References in corpus (2)
Cited by in corpus (38)
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- VideoGPT: Video Generation using VQ-VAE and Transformers
- Convolutional Tensor-Train LSTM for Spatio-temporal Learning
- GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
- Model-Predictive Policy Learning with Uncertainty Regularization for Driving in Dense Traffic
- A review of radar-based nowcasting of precipitation and applicable machine learning techniques
- Symbolic Pregression: Discovering Physical Laws from Distorted Video
- Adjustable Real-time Style Transfer
- Model-Based Visual Planning with Self-Supervised Functional Distances
- Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal Generation
- Experience-Embedded Visual Foresight
- Time Reversal as Self-Supervision
- Video Generation from Single Semantic Label Map
- Preventing Posterior Collapse with delta-VAEs
- Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation
- Effect of Architectures and Training Methods on the Performance of Learned Video Frame Prediction
- Augmenting Physical Simulators with Stochastic Neural Networks: Case Study of Planar Pushing and Bouncing
- Order Matters: Shuffling Sequence Generation for Video Prediction
- Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video Prediction
- Relational Action Forecasting
- Video Interpolation and Prediction with Unsupervised Landmarks
- A Neurally-Inspired Hierarchical Prediction Network for Spatiotemporal Sequence Learning and Prediction
- Precipitation nowcasting using a stochastic variational frame predictor with learned prior distribution
- Joint Training of Variational Auto-Encoder and Latent Energy-Based Model
- Deep Video Prediction for Time Series Forecasting
- Deep Learned Frame Prediction for Video Compression
- Time-Aware and View-Aware Video Rendering for Unsupervised Representation Learning
- Structural Forecasting for Tropical Cyclone Intensity Prediction: Providing Insight with Deep Learning
- Learning Temporal Dynamics from Cycles in Narrated Video
- VUNet: Dynamic Scene View Synthesis for Traversability Estimation using an RGB Camera
- Sound2Sight: Generating Visual Dynamics from Sound and Context
- Hierarchical Video Generation for Complex Data
- High Performance Across Two Atari Paddle Games Using the Same Perceptual Control Architecture Without Training
- VP-GO: a "light" action-conditioned visual prediction model
- Dual-MTGAN: Stochastic and Deterministic Motion Transfer for Image-to-Video Synthesis
- Stochastic Dynamics for Video Infilling
- No Need for Interactions: Robust Model-Based Imitation Learning using Neural ODE
- Learning Semantic-Aware Dynamics for Video Prediction