Learning to Generate Long-term Future via Hierarchical Prediction
arXiv:1704.05831
Abstract
We propose a hierarchical approach for making long-term predictions of future frames. To avoid inherent compounding errors in recursive pixel-level prediction, we propose to first estimate high-level structure in the input frames, then predict how that structure evolves in the future, and finally by observing a single frame from the past and the predicted high-level structure, we construct the future frames without having to observe any of the pixel-level predictions. Long-term video prediction is difficult to perform by recurrently observing the predicted frames because the small errors in pixel space exponentially amplify as predictions are made deeper into the future. Our approach prevents pixel-level error propagation from happening by removing the need to observe the predicted frames. Our model is built with a combination of LSTM and analogy based encoder-decoder convolutional neural networks, which independently predict the video structure and generate the future frames, respectively. In experiments, our model is evaluated on the Human3.6M and Penn Action datasets on the task of long-term pixel-level video prediction of humans performing actions and demonstrate significantly better results than the state-of-the-art.
International Conference on Machine Learning (ICML) 2017
References in corpus (2)
Cited by in corpus (64)
- Overcoming Limitations of Mixture Density Networks: A Sampling and Fitting Framework for Multimodal Future Prediction
- Modeling Human Motion with Quaternion-based Neural Networks
- Adversarial Video Generation on Complex Datasets
- Training Confidence-calibrated Classifiers for Detecting Out-of-Distribution Samples
- Learning to Decompose and Disentangle Representations for Video Prediction
- Learning to dance: A graph convolutional adversarial network to generate realistic dance motions from audio
- Dynamic Facial Expression Generation on Hilbert Hypersphere with Conditional Wasserstein Generative Adversarial Nets
- ConvTransformer: A Convolutional Transformer Network for Video Frame Synthesis
- Inferring Semantic Layout for Hierarchical Text-to-Image Synthesis
- Towards Binary-Valued Gates for Robust LSTM Training
- 3D Human Pose Estimation in the Wild by Adversarial Learning
- Uncertainty Estimation in Autoregressive Structured Prediction
- Future Semantic Segmentation with Convolutional LSTM
- Predicting Video with VQVAE
- Photo-Realistic Video Prediction on Natural Videos of Largely Changing Frames
- PredNet and Predictive Coding: A Critical Review
- Unsupervised Learning of Object Structure and Dynamics from Videos
- Every Smile is Unique: Landmark-Guided Diverse Smile Generation
- Learning to Generate Time-Lapse Videos Using Multi-Stage Dynamic Generative Adversarial Networks
- End-to-End Chinese Landscape Painting Creation Using Generative Adversarial Networks
- Effect of Architectures and Training Methods on the Performance of Learned Video Frame Prediction
- Learning Regularity in Skeleton Trajectories for Anomaly Detection in Videos
- Flow-Grounded Spatial-Temporal Video Prediction from Still Images
- Im2Flow: Motion Hallucination from Static Images for Action Recognition
- Human Motion Transfer from Poses in the Wild
- Disentangling Propagation and Generation for Video Prediction
- DwNet: Dense warp-based network for pose-guided human video generation
- Spatio-Temporal Convolutional LSTMs for Tumor Growth Prediction by Learning 4D Longitudinal Patient Data
- Learning to Forecast and Refine Residual Motion for Image-to-Video Generation
- Forward Prediction for Physical Reasoning
- Relational Action Forecasting
- Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video Prediction
- Synthesizing Long-Term 3D Human Motion and Interaction in 3D Scenes
- A Neurally-Inspired Hierarchical Prediction Network for Spatiotemporal Sequence Learning and Prediction
- FlexLip: A Controllable Text-to-Lip System
- Review of Video Predictive Understanding: Early Action Recognition and Future Action Prediction
- Learning Long-term Visual Dynamics with Region Proposal Interaction Networks
- Deep Learned Frame Prediction for Video Compression
- Better Guider Predicts Future Better: Difference Guided Generative Adversarial Networks
- Hierarchical Model for Long-term Video Prediction
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic Cameras
- Video Time: Properties, Encoders and Evaluation
- Learning to Forecast Videos of Human Activity with Multi-granularity Models and Adaptive Rendering
- Learning Temporal Dynamics from Cycles in Narrated Video
- Recurrent Flow-Guided Semantic Forecasting
- Neural Articulated Radiance Field
- Motion Selective Prediction for Video Frame Synthesis
- HRVGAN: High Resolution Video Generation using Spatio-Temporal GAN
- Predicting the Future with Transformational States
- Inserting Videos into Videos
- Future Video Synthesis with Object Motion Prediction
- Painting Many Pasts: Synthesizing Time Lapse Videos of Paintings
- Long-Term Human Video Generation of Multiple Futures Using Poses
- Efficient training for future video generation based on hierarchical disentangled representation of latent variables
- Taylor saves for later: disentanglement for video prediction using Taylor representation
- A Variational Auto-Encoder Model for Stochastic Point Processes
- Learning Semantic-Aware Dynamics for Video Prediction
- Hierarchical Motion Understanding via Motion Programs
- Layered Controllable Video Generation
- MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics
- SemCo: Toward Semantic Coherent Visual Relationship Forecasting
- Cross-Identity Motion Transfer for Arbitrary Objects through Pose-Attentive Video Reassembling
- Compositional Video Prediction
- Towards Purely Unsupervised Disentanglement of Appearance and Shape for Person Images Generation