Time-Agnostic Prediction: Predicting Predictable Video Frames
arXiv:1808.07784
Abstract
Prediction is arguably one of the most basic functions of an intelligent system. In general, the problem of predicting events in the future or between two waypoints is exceedingly difficult. However, most phenomena naturally pass through relatively predictable bottlenecks---while we cannot predict the precise trajectory of a robot arm between being at rest and holding an object up, we can be certain that it must have picked the object up. To exploit this, we decouple visual prediction from a rigid notion of time. While conventional approaches predict frames at regularly spaced temporal intervals, our time-agnostic predictors (TAP) are not tied to specific times so that they may instead discover predictable "bottleneck" frames no matter when they occur. We evaluate our approach for future and intermediate frame prediction across three robotic manipulation tasks. Our predictions are not only of higher visual quality, but also correspond to coherent semantic subgoals in temporally extended tasks.
8 pages, plus appendices
References in corpus (4)
Cited by in corpus (14)
- Overcoming Limitations of Mixture Density Networks: A Sampling and Fitting Framework for Multimodal Future Prediction
- Rapid trial-and-error learning with simulation supports flexible tool use and physical reasoning
- Procedure Planning in Instructional Videos
- Adaptive Skip Intervals: Temporal Abstraction for Recurrent Dynamical Models
- Vid2Param: Modelling of Dynamics Parameters from Video
- Identifying Critical States by the Action-Based Variance of Expected Return
- Temporal Dynamic Model for Resting State fMRI Data: A Neural Ordinary Differential Equation approach
- Beta DVBF: Learning State-Space Models for Control from High Dimensional Observations
- Causal Future Prediction in a Minkowski Space-Time
- Variational Predictive Routing with Nested Subjective Timescales
- On the potential for open-endedness in neural networks
- Point-to-Point Video Generation
- Stochastic Dynamics for Video Infilling
- Generating the support with extreme value losses