DESIRE: Distant Future Prediction in Dynamic Scenes with Interacting Agents
arXiv:1704.04394
Abstract
We introduce a Deep Stochastic IOC RNN Encoderdecoder framework, DESIRE, for the task of future predictions of multiple interacting agents in dynamic scenes. DESIRE effectively predicts future locations of objects in multiple scenes by 1) accounting for the multi-modal nature of the future prediction (i.e., given the same context, future may vary), 2) foreseeing the potential future outcomes and make a strategic prediction based on that, and 3) reasoning not only from the past motion history, but also from the scene context as well as the interactions among the agents. DESIRE achieves these in a single end-to-end trainable neural network model, while being computationally efficient. The model first obtains a diverse set of hypothetical future prediction samples employing a conditional variational autoencoder, which are ranked and refined by the following RNN scoring-regression module. Samples are scored by accounting for accumulated future rewards, which enables better long-term strategic decisions similar to IOC frameworks. An RNN scene context fusion module jointly captures past motion histories, the semantic scene context and interactions among multiple agents. A feedback mechanism iterates over the ranking and refinement to further boost the prediction accuracy. We evaluate our model on two publicly available datasets: KITTI and Stanford Drone Dataset. Our experiments show that the proposed model significantly improves the prediction accuracy compared to other baseline methods.
Accepted at CVPR 2017
References in corpus (3)
Cited by in corpus (29)
- Argoverse: 3D Tracking and Forecasting with Rich Maps
- Social-BiGAT: Multimodal Trajectory Forecasting using Bicycle-GAN and Graph Attention Networks
- An Evaluation of Trajectory Prediction Approaches and Notes on the TrajNet Benchmark
- THOMAS: Trajectory Heatmap Output with learned Multi-Agent Sampling
- Conditional Flow Variational Autoencoders for Structured Sequence Prediction
- Deep Imitative Models for Flexible Inference, Planning, and Control
- Situation-Aware Pedestrian Trajectory Prediction with Spatio-Temporal Attention Model
- Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks
- Accurate and Diverse Sampling of Sequences based on a "Best of Many" Sample Objective
- Trajformer: Trajectory Prediction with Local Self-Attentive Contexts for Autonomous Driving
- Heterogeneous Edge-Enhanced Graph Attention Network For Multi-Agent Trajectory Prediction
- Value Propagation Networks
- A New Multi-vehicle Trajectory Generator to Simulate Vehicle-to-Vehicle Encounters
- BiTraP: Bi-directional Pedestrian Trajectory Prediction with Multi-modal Goal Estimation
- TPNet: Trajectory Proposal Network for Motion Prediction
- Im2Flow: Motion Hallucination from Static Images for Action Recognition
- Multi-Agent Reinforcement Learning with Multi-Step Generative Models
- CAR-Net: Clairvoyant Attentive Recurrent Network
- Relational Action Forecasting
- Convolutional Neural Network for Trajectory Prediction
- Naturalistic Driver Intention and Path Prediction using Recurrent Neural Networks
- DiversityGAN: Diversity-Aware Vehicle Motion Prediction via Latent Semantic Sampling
- Understanding Human Behaviors in Crowds by Imitating the Decision-Making Process
- Future Person Localization in First-Person Videos
- HGCN-GJS: Hierarchical Graph Convolutional Network with Groupwise Joint Sampling for Trajectory Prediction
- AVGCN: Trajectory Prediction using Graph Convolutional Networks Guided by Human Attention
- Deep Structured Reactive Planning
- Improved Activity Forecasting for Generating Trajectories
- Euro-PVI: Pedestrian Vehicle Interactions in Dense Urban Centers