Generating Videos with Scene Dynamics
arXiv:1609.02612
Abstract
We capitalize on large amounts of unlabeled video in order to learn a model of scene dynamics for both video recognition tasks (e.g. action classification) and video generation tasks (e.g. future prediction). We propose a generative adversarial network for video with a spatio-temporal convolutional architecture that untangles the scene's foreground from the background. Experiments suggest this model can generate tiny videos up to a second at full frame rate better than simple baselines, and we show its utility at predicting plausible futures of static images. Moreover, experiments and visualizations show the model internally learns useful features for recognizing actions with minimal supervision, suggesting scene dynamics are a promising signal for representation learning. We believe generative video models can impact many applications in video understanding and simulation.
NIPS 2016. See more at http://web.mit.edu/vondrick/tinyvideo/
References in corpus (1)
Cited by in corpus (34)
- Deep Learning for Precipitation Nowcasting: A Benchmark and A New Model
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- Learning to Generate Images of Outdoor Scenes from Attributes and Semantic Layouts
- Self-Supervised Visual Planning with Temporal Skip Connections
- Photographic Image Synthesis with Cascaded Refinement Networks
- Lat-Net: Compressing Lattice Boltzmann Flow Simulations using Deep Neural Networks
- Transformation-Grounded Image Generation Network for Novel 3D View Synthesis
- Generative Adversarial Networks for Electronic Health Records: A Framework for Exploring and Evaluating Methods for Predicting Drug-Induced Laboratory Test Trajectories
- Conditional Adversarial Network for Semantic Segmentation of Brain Tumor
- Unsupervised Representation Learning by Sorting Sequences
- Video Imagination from a Single Image with Transformation Generation
- DiscrimNet: Semi-Supervised Action Recognition from Videos using Generative Adversarial Networks
- Message Passing Multi-Agent GANs
- Dual Motion GAN for Future-Flow Embedded Video Prediction
- Dynamics Transfer GAN: Generating Video by Transferring Arbitrary Temporal Dynamics from a Source Video to a Single Target Image
- Distributional Adversarial Networks
- Vid2Game: Controllable Characters Extracted from Real-World Videos
- Face Translation between Images and Videos using Identity-aware CycleGAN
- Planning Robot Motion using Deep Visual Prediction
- Relational Action Forecasting
- Visual Forecasting by Imitating Dynamics in Natural Sequences
- Hybrid VAE: Improving Deep Generative Models using Partial Observations
- Predictive Learning: Using Future Representation Learning Variantial Autoencoder for Human Action Prediction
- Deep Learning for Low-Dose CT Denoising
- Image2GIF: Generating Cinemagraphs using Recurrent Deep Q-Networks
- Disentangling Motion, Foreground and Background Features in Videos
- Human Pose Forecasting via Deep Markov Models
- Sliced Wasserstein Generative Models
- Hierarchical Model for Long-term Video Prediction
- Classification of sparsely labeled spatio-temporal data through semi-supervised adversarial learning
- Visual Data Augmentation through Learning
- Concept Formation and Dynamics of Repeated Inference in Deep Generative Models
- Automatic Realistic Music Video Generation from Segments of Youtube Videos
- cvpaper.challenge in 2016: Futuristic Computer Vision through 1,600 Papers Survey