Accurate and Diverse Sampling of Sequences based on a "Best of Many" Sample Objective
arXiv:1806.07772
Abstract
For autonomous agents to successfully operate in the real world, anticipation of future events and states of their environment is a key competence. This problem has been formalized as a sequence extrapolation problem, where a number of observations are used to predict the sequence into the future. Real-world scenarios demand a model of uncertainty of such predictions, as predictions become increasingly uncertain -- in particular on long time horizons. While impressive results have been shown on point estimates, scenarios that induce multi-modal distributions over future sequences remain challenging. Our work addresses these challenges in a Gaussian Latent Variable model for sequence prediction. Our core contribution is a "Best of Many" sample objective that leads to more accurate and more diverse predictions that better capture the true variations in real-world sequence data. Beyond our analysis of improved model fit, our models also empirically outperform prior work on three diverse tasks ranging from traffic scenes to weather data.
Added additional references and baselines. (Appeared in CVPR 2018)
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Deep Learning for Precipitation Nowcasting: A Benchmark and A New Model
- Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks
- DESIRE: Distant Future Prediction in Dynamic Scenes with Interacting Agents
- Techniques for Learning Binary Stochastic Feedforward Neural Networks
- Motion Prediction Under Multimodality with Conditional Stochastic Networks
- Simplified Stochastic Feedforward Neural Networks
Cited by in corpus (12)
- GAN-Leaks: A Taxonomy of Membership Inference Attacks against Generative Models
- Machine Learning for Spatiotemporal Sequence Forecasting: A Survey
- Conditional Flow Variational Autoencoders for Structured Sequence Prediction
- Diverse Human Motion Prediction Guided by Multi-Level Spatial-Temporal Anchors
- Unsupervised Learning of Object Structure and Dynamics from Videos
- BiTraP: Bi-directional Pedestrian Trajectory Prediction with Multi-modal Goal Estimation
- Diverse Multimedia Layout Generation with Multi Choice Learning
- We are More than Our Joints: Predicting how 3D Bodies Move
- Motion Prediction using Trajectory Sets and Self-Driving Domain Knowledge
- Back to square one: probabilistic trajectory forecasting without bells and whistles
- Learning from Demonstration with Weakly Supervised Disentanglement
- Euro-PVI: Pedestrian Vehicle Interactions in Dense Urban Centers