Deep End2End Voxel2Voxel Prediction
arXiv:1511.06681
Abstract
Over the last few years deep learning methods have emerged as one of the most prominent approaches for video analysis. However, so far their most successful applications have been in the area of video classification and detection, i.e., problems involving the prediction of a single class label or a handful of output variables per video. Furthermore, while deep networks are commonly recognized as the best models to use in these domains, there is a widespread perception that in order to yield successful results they often require time-consuming architecture search, manual tweaking of parameters and computationally intensive pre-processing or post-processing methods. In this paper we challenge these views by presenting a deep 3D convolutional architecture trained end to end to perform voxel-level prediction, i.e., to output a variable at every voxel of the video. Most importantly, we show that the same exact architecture can be used to achieve competitive results on three widely different voxel-prediction tasks: video semantic segmentation, optical flow estimation, and video coloring. The three networks learned on these problems are trained from raw video without any form of preprocessing and their outputs do not require post-processing to achieve outstanding performance. Thus, they offer an efficient alternative to traditional and much more computationally expensive methods in these video domains.
References in corpus (7)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Going Deeper with Convolutions
- Fully Convolutional Networks for Semantic Segmentation
- Learning Spatiotemporal Features with 3D Convolutional Networks
- Geometric Context from Videos
Cited by in corpus (8)
- 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation
- VoxResNet: Deep Voxelwise Residual Networks for Volumetric Brain Segmentation
- Flow-Guided Feature Aggregation for Video Object Detection
- Temporal Activity Detection in Untrimmed Videos with Recurrent Neural Networks
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- Hybrid Learning of Optical Flow and Next Frame Prediction to Boost Optical Flow in the Wild
- Video Frame Interpolation by Plug-and-Play Deep Locally Linear Embedding
- Generalized Deep Image to Image Regression