Deep multi-scale video prediction beyond mean square error
arXiv:1511.05440
Abstract
Learning to predict future images from a video sequence involves the construction of an internal representation that models the image evolution accurately, and therefore, to some degree, its content and dynamics. This is why pixel-space video prediction may be viewed as a promising avenue for unsupervised feature learning. In addition, while optical flow has been a very studied problem in computer vision for a long time, future frame prediction is rarely approached. Still, many vision applications could benefit from the knowledge of the next frames of videos, that does not require the complexity of tracking every pixel trajectories. In this work, we train a convolutional network to generate future frames given an input sequence. To deal with the inherently blurry predictions obtained from the standard Mean Squared Error (MSE) loss function, we propose three different and complementary feature learning strategies: a multi-scale architecture, an adversarial training method, and an image gradient difference loss function. We compare our predictions to different published results based on recurrent neural networks on the UCF101 dataset
References in corpus (4)
Cited by in corpus (88)
- Video Enhancement with Task-Oriented Flow
- A trans-disciplinary review of deep learning research for water resources scientists
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
- Autoencoding beyond pixels using a learned similarity metric
- Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
- Sharpness-aware Low dose CT denoising using conditional generative adversarial network
- Deep Learning for Physical Processes: Incorporating Prior Scientific Knowledge
- Disentangling factors of variation in deep representations using adversarial training
- Mode Regularized Generative Adversarial Networks
- Overcoming Limitations of Mixture Density Networks: A Sampling and Fitting Framework for Multimodal Future Prediction
- Learning Depth from Single Images with Deep Neural Network Embedding Focal Length
- Learning a Driving Simulator
- On Learning 3D Face Morphable Model from In-the-wild Images
- Modeling Human Motion with Quaternion-based Neural Networks
- Adversarial Video Generation on Complex Datasets
- Space-Time Correspondence as a Contrastive Random Walk
- Train Sparsely, Generate Densely: Memory-efficient Unsupervised Training of High-resolution Temporal GAN
- Controlling generative models with continuous factors of variations
- Collaborative Learning for Faster StyleGAN Embedding
- Cross-view image synthesis using geometry-guided conditional GANs
- Adversarial PoseNet: A Structure-aware Convolutional Network for Human Pose Estimation
- It GAN DO Better: GAN-based Detection of Objects on Images with Varying Quality
- Effective Image Differencing with ConvNets for Real-time Transient Hunting
- Transformation-Grounded Image Generation Network for Novel 3D View Synthesis
- Generative Adversarial Networks for Electronic Health Records: A Framework for Exploring and Evaluating Methods for Predicting Drug-Induced Laboratory Test Trajectories
- Multi-Modal Imitation Learning from Unstructured Demonstrations using Generative Adversarial Nets
- Comparing recurrent and convolutional neural networks for predicting wave propagation
- Multi-view Generative Adversarial Networks
- Conditional Adversarial Network for Semantic Segmentation of Brain Tumor
- High Perceptual Quality Image Denoising with a Posterior Sampling CGAN
- Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training
- STARDATA: A StarCraft AI Research Dataset
- Medical Image Synthesis with Context-Aware Generative Adversarial Networks
- Time-Agnostic Prediction: Predicting Predictable Video Frames
- Optimal Physical Preprocessing for Example-Based Super-Resolution
- DiscrimNet: Semi-Supervised Action Recognition from Videos using Generative Adversarial Networks
- A multiscale neural network based on hierarchical matrices
- Mechanisms of a Convolutional Neural Network for Learning Three-dimensional Unsteady Wake Flow
- CariGAN: Caricature Generation through Weakly Paired Adversarial Learning
- Dynamics Transfer GAN: Generating Video by Transferring Arbitrary Temporal Dynamics from a Source Video to a Single Target Image
- DOOM Level Generation using Generative Adversarial Networks
- Unsupervised Video-to-Video Translation
- Self-Diagnosing GAN: Diagnosing Underrepresented Samples in Generative Adversarial Networks
- Hybrid Learning of Optical Flow and Next Frame Prediction to Boost Optical Flow in the Wild
- Towards High-fidelity Nonlinear 3D Face Morphable Model
- Learning to Predict Without Looking Ahead: World Models Without Forward Prediction
- Improving GAN Training via Binarized Representation Entropy (BRE) Regularization
- Self-Supervised Relative Depth Learning for Urban Scene Understanding
- Order Matters: Shuffling Sequence Generation for Video Prediction
- Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video Prediction
- Deep learning approach in multi-scale prediction of turbulent mixing-layer
- Learning Correspondence from the Cycle-Consistency of Time
- Deep Variational Inference Without Pixel-Wise Reconstruction
- Learning Long-term Visual Dynamics with Region Proposal Interaction Networks
- Learning Temporal Transformations From Time-Lapse Videos
- Very Long Natural Scenery Image Prediction by Outpainting
- Predicting Real-Time Locational Marginal Prices: A GAN-Based Video Prediction Approach
- Image2GIF: Generating Cinemagraphs using Recurrent Deep Q-Networks
- Video Extrapolation with an Invertible Linear Embedding
- Generative Image Modeling using Style and Structure Adversarial Networks
- Neural Allocentric Intuitive Physics Prediction from Real Videos
- Adversarial Feature Desensitization
- Learning Temporal Dynamics from Cycles in Narrated Video
- Diverse Video Generation using a Gaussian Process Trigger
- Learning to Forecast Videos of Human Activity with Multi-granularity Models and Adaptive Rendering
- On the difficulty of learning and predicting the long-term dynamics of bouncing objects
- Position-Aware Convolutional Networks for Traffic Prediction
- PDWN: Pyramid Deformable Warping Network for Video Interpolation
- Interpretable Intuitive Physics Model
- Hierarchical Video Generation for Complex Data
- Pairwise-GAN: Pose-based View Synthesis through Pair-Wise Training
- Spatio-Temporal Image Boundary Extrapolation
- Learning Dynamical Systems from Noisy Sensor Measurements using Multiple Shooting
- Progressive VAE Training on Highly Sparse and Imbalanced Data
- Complex Valued Gated Auto-encoder for Video Frame Prediction
- Inserting Videos into Videos
- Linear Discriminant Generative Adversarial Networks
- Non-destructive three-dimensional measurement of hand vein based on self-supervised network
- Cross-Modality Distillation: A case for Conditional Generative Adversarial Networks
- Unbiased Image Style Transfer
- Generative Actor-Critic: An Off-policy Algorithm Using the Push-forward Model
- Pipeline for 3D reconstruction of the human body from AR/VR headset mounted egocentric cameras
- Location Dependency in Video Prediction
- Real-time Locational Marginal Price Forecasting Using Generative Adversarial Network
- DINO: A Conditional Energy-Based GAN for Domain Translation
- Fine-scale Surface Normal Estimation using a Single NIR Image
- Multimodal Image Synthesis with Conditional Implicit Maximum Likelihood Estimation
- Stochastic Dynamics for Video Infilling