Predicting Deeper into the Future of Semantic Segmentation
arXiv:1703.07684
Abstract
The ability to predict and therefore to anticipate the future is an important attribute of intelligence. It is also of utmost importance in real-time systems, e.g. in robotics or autonomous driving, which depend on visual scene understanding for decision making. While prediction of the raw RGB pixel values in future video frames has been studied in previous work, here we introduce the novel task of predicting semantic segmentations of future frames. Given a sequence of video frames, our goal is to predict segmentation maps of not yet observed video frames that lie up to a second or further in the future. We develop an autoregressive convolutional neural network that learns to iteratively generate multiple frames. Our results on the Cityscapes dataset show that directly predicting future segmentations is substantially better than predicting and then segmenting future RGB frames. Prediction results up to half a second in the future are visually convincing and are much more accurate than those of a baseline based on warping semantic segmentations using optical flow.
Accepted to ICCV 2017. Supplementary material available on the authors' webpages
References in corpus (8)
- Distilling the Knowledge in a Neural Network
- Generating Videos with Scene Dynamics
- Semantic Segmentation using Adversarial Networks
- LR-GAN: Layered Recursive Generative Adversarial Networks for Image Generation
- Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks
- Transformation-Based Models of Video Sequences
- Unsupervised Learning of Long-Term Motion Dynamics for Videos
- Video Scene Parsing with Predictive Feature Learning
Cited by in corpus (19)
- Overcoming Limitations of Mixture Density Networks: A Sampling and Fitting Framework for Multimodal Future Prediction
- Convolutional Invasion and Expansion Networks for Tumor Growth Prediction
- Predicting Video with VQVAE
- Novel Video Prediction for Large-scale Scene using Optical Flow
- CortexNet: a Generic Network Family for Robust Visual Temporal Representations
- Learning to Represent Mechanics via Long-term Extrapolation and Interpolation
- Revisiting Hierarchical Approach for Persistent Long-Term Video Prediction
- Learning Representations for Predicting Future Activities
- Relational Action Forecasting
- A Novel Stochastic Stratified Average Gradient Method: Convergence Rate and Its Complexity
- ClusterFit: Improving Generalization of Visual Representations
- Semantic Road Layout Understanding by Generative Adversarial Inpainting
- Future Segmentation Using 3D Structure
- Revisiting Deep Architectures for Head Motion Prediction in 360° Videos
- Bayesian Prediction of Future Street Scenes through Importance Sampling based Optimization
- MLPerf Mobile Inference Benchmark
- Recurrent Flow-Guided Semantic Forecasting
- Future Video Synthesis with Object Motion Prediction
- Long-Term On-Board Prediction of People in Traffic Scenes under Uncertainty