Video (language) modeling: a baseline for generative models of natural videos
arXiv:1412.6604
Abstract
We propose a strong baseline model for unsupervised feature learning using video data. By learning to predict missing frames or extrapolate future frames from an input video sequence, the model discovers both spatial and temporal correlations which are useful to represent complex deformations and motion patterns. The models we propose are largely borrowed from the language modeling literature, and adapted to the vision domain by quantizing the space of image patches into a large dictionary. We demonstrate the approach on both a filling and a generation task. For the first time, we show that, after training on natural videos, such a model can predict non-trivial motions over short video sequences.
References in corpus (5)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Simultaneous Detection and Segmentation
- Fast Inference in Sparse Coding Algorithms with Applications to Object Recognition
- Emergence of Complex-Like Cells in a Temporal Product Network with Local Receptive Fields
Cited by in corpus (120)
- Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
- Unsupervised Learning of Video Representations using LSTMs
- Deep Learning for Precipitation Nowcasting: A Benchmark and A New Model
- Image and Video Compression with Neural Networks: A Review
- Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
- A Review on Deep Learning Techniques for Video Prediction
- Unsupervised Learning for Physical Interaction through Video Prediction
- Stochastic Adversarial Video Prediction
- Describing Videos by Exploiting Temporal Structure
- Adversarial Video Generation on Complex Datasets
- Train Sparsely, Generate Densely: Memory-efficient Unsupervised Training of High-resolution Temporal GAN
- Learning to Decompose and Disentangle Representations for Video Prediction
- Going in circles is the way forward: the role of recurrence in visual inference
- Few-Shot Deep Adversarial Learning for Video-based Person Re-identification
- Generative Image Modeling Using Spatial LSTMs
- Unsupervised Learning of View-invariant Action Representations
- Machine Learning for Spatiotemporal Sequence Forecasting: A Survey
- Video Frame Interpolation via Adaptive Separable Convolution
- Stochastic Variational Video Prediction
- Convolutional Invasion and Expansion Networks for Tumor Growth Prediction
- Recurrent Network Models for Human Dynamics
- Transformation-Based Models of Video Sequences
- Learning to Perform Physics Experiments via Deep Reinforcement Learning
- Deep Recurrent Convolutional Networks for Video-based Person Re-identification: An End-to-End Approach
- Symbolic Pregression: Discovering Physical Laws from Distorted Video
- Colorization as a Proxy Task for Visual Understanding
- Structured Sequence Modeling with Graph Convolutional Recurrent Networks
- YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video
- Sea Level Anomaly Prediction using Recurrent Neural Networks
- Temporal Generative Adversarial Nets with Singular Value Clipping
- Predicting Deeper into the Future of Semantic Segmentation
- Video Frame Interpolation via Adaptive Convolution
- Learning to Linearize Under Uncertainty
- Unsupervised Learning of Long-Term Motion Dynamics for Videos
- Future Semantic Segmentation with Convolutional LSTM
- Adversarial Learning with Local Coordinate Coding
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation
- Context-aware Synthesis for Video Frame Interpolation
- FitVid: Overfitting in Pixel-Level Video Prediction
- HybridNet: Integrating Model-based and Data-driven Learning to Predict Evolution of Dynamical Systems
- Video Imagination from a Single Image with Transformation Generation
- Modeling Spatial-Temporal Clues in a Hybrid Deep Learning Framework for Video Classification
- Predicting Video with VQVAE
- Time-Agnostic Prediction: Predicting Predictable Video Frames
- Photo-Realistic Video Prediction on Natural Videos of Largely Changing Frames
- Memory In Memory: A Predictive Neural Network for Learning Higher-Order Non-Stationarity from Spatiotemporal Dynamics
- CloudCast: A Satellite-Based Dataset and Baseline for Forecasting Clouds
- High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks
- A neural network trained to predict future video frames mimics critical properties of biological neuronal responses and perception
- Dense Optical Flow Prediction from a Static Image
- Unsupervised Learning of Object Structure and Dynamics from Videos
- Unsupervised Learning from Continuous Video in a Scalable Predictive Recurrent Network
- Video Generation from Single Semantic Label Map
- Improving Video Generation for Multi-functional Applications
- Transformation-based Adversarial Video Prediction on Large-Scale Data
- Novel Video Prediction for Large-scale Scene using Optical Flow
- Mutual Suppression Network for Video Prediction using Disentangled Features
- Im2Flow: Motion Hallucination from Static Images for Action Recognition
- Disentangling Propagation and Generation for Video Prediction
- Procedure Planning in Instructional Videos
- Answering Visual What-If Questions: From Actions to Predicted Scene Descriptions
- MotionRNN: A Flexible Model for Video Prediction with Spacetime-Varying Motions
- Interactive Fusion of Multi-level Features for Compositional Activity Recognition
- Revisiting Hierarchical Approach for Persistent Long-Term Video Prediction
- Self-Supervised Relative Depth Learning for Urban Scene Understanding
- Relational Action Forecasting
- Learning Representations for Predicting Future Activities
- Unsupervised Learning of Edges
- CUPID: Adaptive Curation of Pre-training Data for Video-and-Language Representation Learning
- Visual Language Modeling on CNN Image Representations
- Adversarial Video Compression Guided by Soft Edge Detection
- The Best of Both Worlds: Combining Data-independent and Data-driven Approaches for Action Recognition
- Building Machines that Learn and Think for Themselves: Commentary on Lake et al., Behavioral and Brain Sciences, 2017
- Understanding the Tradeoffs in Client-side Privacy for Downstream Speech Tasks
- An Uncertain Future: Forecasting from Static Images using Variational Autoencoders
- FDNet: A Deep Learning Approach with Two Parallel Cross Encoding Pathways for Precipitation Nowcasting
- Multi-View Frame Reconstruction with Conditional GAN
- Review of Video Predictive Understanding: Early Action Recognition and Future Action Prediction
- Learning Temporal Transformations From Time-Lapse Videos
- Future Segmentation Using 3D Structure
- Deep Video Prediction for Time Series Forecasting
- Future Frame Prediction of a Video Sequence
- ContextVP: Fully Context-Aware Video Prediction
- Learning image representations tied to ego-motion
- Long-Term Image Boundary Prediction
- Learning Temporal Embeddings for Complex Video Analysis
- A Temporally-Aware Interpolation Network for Video Frame Inpainting
- Animating Landscape: Self-Supervised Learning of Decoupled Motion and Appearance for Single-Image Video Synthesis
- Image2GIF: Generating Cinemagraphs using Recurrent Deep Q-Networks
- Video Time: Properties, Encoders and Evaluation
- Auto-Embedding Generative Adversarial Networks for High Resolution Image Synthesis
- Latent Neural Differential Equations for Video Generation
- Learning Temporal Dynamics from Cycles in Narrated Video
- On the difficulty of learning and predicting the long-term dynamics of bouncing objects
- Learning to Forecast Videos of Human Activity with Multi-granularity Models and Adaptive Rendering
- Wide and Narrow: Video Prediction from Context and Motion
- Temporal Dynamic Model for Resting State fMRI Data: A Neural Ordinary Differential Equation approach
- Position-based Content Attention for Time Series Forecasting with Sequence-to-sequence RNNs
- The LICORS Cabinet: Nonparametric Algorithms for Spatio-temporal Prediction
- Painting Many Pasts: Synthesizing Time Lapse Videos of Paintings
- Recurrent Flow-Guided Semantic Forecasting
- Hierarchical Video Generation for Complex Data
- ModeRNN: Harnessing Spatiotemporal Mode Collapse in Unsupervised Predictive Learning
- Sound2Sight: Generating Visual Dynamics from Sound and Context
- One-Step Time-Dependent Future Video Frame Prediction with a Convolutional Encoder-Decoder Neural Network
- Spatio-Temporal Image Boundary Extrapolation
- Learning Image Matching by Simply Watching Video
- Improving Generative Adversarial Networks with Local Coordinate Coding
- Action Anticipation with RBF Kernelized Feature Mapping RNN
- Understanding in Artificial Intelligence
- A Framework for Multisensory Foresight for Embodied Agents
- Unsupervised Learning Layers for Video Analysis
- Point-to-Point Video Generation
- From Single to Multiple: Leveraging Multi-level Prediction Spaces for Video Forecasting
- Multi-Modal Temporal Convolutional Network for Anticipating Actions in Egocentric Videos
- Unsupervised Video Prediction from a Single Frame by Estimating 3D Dynamic Scene Structure
- From Thumbnails to Summaries - A single Deep Neural Network to Rule Them All
- Action-conditional Sequence Modeling for Recommendation
- SDCNet: Video Prediction Using Spatially-Displaced Convolution
- Cubic LSTMs for Video Prediction