Unsupervised Keypoint Learning for Guiding Class-Conditional Video Prediction
arXiv:1910.02027
Abstract
We propose a deep video prediction model conditioned on a single image and an action class. To generate future frames, we first detect keypoints of a moving object and predict future motion as a sequence of keypoints. The input image is then translated following the predicted keypoints sequence to compose future frames. Detecting the keypoints is central to our algorithm, and our method is trained to detect the keypoints of arbitrary objects in an unsupervised manner. Moreover, the detected keypoints of the original videos are used as pseudo-labels to learn the motion of objects. Experimental results show that our method is successfully applied to various datasets without the cost of labeling keypoints in videos. The detected keypoints are similar to human-annotated labels, and prediction results are more realistic compared to the previous methods.
NeurIPS 2019
Cited by in corpus (11)
- Deep Learning for Vision-based Prediction: A Survey
- Human Motion Transfer from Poses in the Wild
- Revisiting Hierarchical Approach for Persistent Long-Term Video Prediction
- Flow Guided Transformable Bottleneck Networks for Motion Retargeting
- PriorityCut: Occlusion-guided Regularization for Warp-based Image Animation
- Future Frame Prediction for Robot-assisted Surgery
- GaussiGAN: Controllable Image Synthesis with 3D Gaussians from Unposed Silhouettes
- AutoTrajectory: Label-free Trajectory Extraction and Prediction from Videos using Dynamic Points
- From Single to Multiple: Leveraging Multi-level Prediction Spaces for Video Forecasting
- Action2video: Generating Videos of Human 3D Actions
- Adaptive Future Frame Prediction with Ensemble Network