Forecasting Human Dynamics from Static Images
arXiv:1704.03432
Abstract
This paper presents the first study on forecasting human dynamics from static images. The problem is to input a single RGB image and generate a sequence of upcoming human body poses in 3D. To address the problem, we propose the 3D Pose Forecasting Network (3D-PFNet). Our 3D-PFNet integrates recent advances on single-image human pose estimation and sequence prediction, and converts the 2D predictions into 3D space. We train our 3D-PFNet using a three-step training strategy to leverage a diverse source of training data, including image and video based human pose datasets and 3D motion capture (MoCap) data. We demonstrate competitive performance of our 3D-PFNet on 2D pose forecasting and 3D pose recovery through quantitative and qualitative results.
Accepted in CVPR 2017
Cited by in corpus (10)
- Learning to Generate Long-term Future via Hierarchical Prediction
- Dance Revolution: Long-Term Dance Generation with Music via Curriculum Learning
- Improving Video Generation for Multi-functional Applications
- Im2Flow: Motion Hallucination from Static Images for Action Recognition
- Flow-Grounded Spatial-Temporal Video Prediction from Still Images
- Learning to Forecast and Refine Residual Motion for Image-to-Video Generation
- Imitation Learning for Human Pose Prediction
- Learning to Predict Diverse Human Motions from a Single Image via Mixture Density Networks
- Recurrent Flow-Guided Semantic Forecasting
- MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics