Dense Optical Flow Prediction from a Static Image
arXiv:1505.00295
Abstract
Given a scene, what is going to move, and in what direction will it move? Such a question could be considered a non-semantic form of action prediction. In this work, we present a convolutional neural network (CNN) based approach for motion prediction. Given a static image, this CNN predicts the future motion of each and every pixel in the image in terms of optical flow. Our CNN model leverages the data in tens of thousands of realistic videos to train our model. Our method relies on absolutely no human labeling and is able to predict motion based on the context of the scene. Because our CNN model makes no assumptions about the underlying scene, it can predict future optical flow on a diverse set of scenarios. We outperform all previous approaches by large margins.
References in corpus (6)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Going Deeper with Convolutions
- Video (language) modeling: a baseline for generative models of natural videos
- Designing Deep Networks for Surface Normal Estimation
Cited by in corpus (4)
- ReconNet: Non-Iterative Reconstruction of Images from Compressively Sensed Random Measurements
- Spatiotemporal Recurrent Convolutional Networks for Traffic Prediction in Transportation Networks
- cvpaper.challenge in 2016: Futuristic Computer Vision through 1,600 Papers Survey
- VP-GO: a "light" action-conditioned visual prediction model