Deep Learning for Vision-based Prediction: A Survey
arXiv:2007.00095
Abstract
Vision-based prediction algorithms have a wide range of applications including autonomous driving, surveillance, human-robot interaction, weather prediction. The objective of this paper is to provide an overview of the field in the past five years with a particular focus on deep learning approaches. For this purpose, we categorize these algorithms into video prediction, action prediction, trajectory prediction, body motion prediction, and other prediction applications. For each category, we highlight the common architectures, training methods and types of data used. In addition, we discuss the common evaluation metrics and datasets used for vision-based prediction tasks. A database of all the information presented in this survey including, cross-referenced according to papers, datasets and metrics, can be found online at https://github.com/aras62/vision-based-prediction.
References in corpus (8)
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- YouTube-8M: A Large-Scale Video Classification Benchmark
- MOTChallenge 2015: Towards a Benchmark for Multi-Target Tracking
- INTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps
- PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding
- Unsupervised Keypoint Learning for Guiding Class-Conditional Video Prediction
- Online Vehicle Trajectory Prediction using Policy Anticipation Network and Optimization-based Context Reasoning
- Dynamic Hilbert Maps: Real-Time Occupancy Predictions in Changing Environment