Learning to Predict Diverse Human Motions from a Single Image via Mixture Density Networks
arXiv:2109.05776
Abstract
Human motion prediction, which plays a key role in computer vision, generally requires a past motion sequence as input. However, in real applications, a complete and correct past motion sequence can be too expensive to achieve. In this paper, we propose a novel approach to predicting future human motions from a much weaker condition, i.e., a single image, with mixture density networks (MDN) modeling. Contrary to most existing deep human motion prediction approaches, the multimodal nature of MDN enables the generation of diverse future motion hypotheses, which well compensates for the strong stochastic ambiguity aggregated by the single input and human motion uncertainty. In designing the loss function, we further introduce the energy-based formulation to flexibly impose prior losses over the learnable parameters of MDN to maintain motion coherence as well as improve the prediction accuracy by customizing the energy functions. Our trained model directly takes an image as input and generates multiple plausible motions that satisfy the given condition. Extensive experiments on two standard benchmark datasets demonstrate the effectiveness of our method in terms of prediction diversity and accuracy.
References in corpus (11)
- Overcoming Limitations of Mixture Density Networks: A Sampling and Fitting Framework for Multimodal Future Prediction
- A simple yet effective baseline for 3d human pose estimation
- Bio-LSTM: A Biomechanically Inspired Recurrent Neural Network for 3D Pedestrian Pose and Gait Prediction
- Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach
- Compositional Human Pose Regression
- Convolutional Sequence to Sequence Model for Human Dynamics
- Generating Multiple Hypotheses for 3D Human Pose Estimation with Mixture Density Network
- Forecasting Human Dynamics from Static Images
- Occlusion-aware Hand Pose Estimation Using Hierarchical Mixture Density Network
- Learning 3D Human Dynamics from Video
- Imitation Learning for Human Pose Prediction