Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video
arXiv:1511.09439
Abstract
This paper addresses the challenge of 3D full-body human pose estimation from a monocular image sequence. Here, two cases are considered: (i) the image locations of the human joints are provided and (ii) the image locations of joints are unknown. In the former case, a novel approach is introduced that integrates a sparsity-driven 3D geometric prior and temporal smoothness. In the latter case, the former case is extended by treating the image locations of the joints as latent variables. A deep fully convolutional network is trained to predict the uncertainty maps of the 2D joint locations. The 3D pose estimates are realized via an Expectation-Maximization algorithm over the entire sequence, where it is shown that the 2D joint location uncertainties can be conveniently marginalized out during inference. Empirical evaluation on the Human3.6M dataset shows that the proposed approaches achieve greater 3D pose estimation accuracy over state-of-the-art baselines. Further, the proposed approach outperforms a publicly available 2D pose estimation baseline on the challenging PennAction dataset.
Published in CVPR2016
References in corpus (5)
Cited by in corpus (21)
- VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera
- Learning from Synthetic Humans
- Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image
- A simple yet effective baseline for 3d human pose estimation
- Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach
- Synthesizing Training Images for Boosting Human 3D Pose Estimation
- Fast and Robust Multi-Person 3D Pose Estimation from Multiple Views
- 3D Human Pose Estimation in the Wild by Adversarial Learning
- Compact Model Representation for 3D Reconstruction
- Learning to Estimate 3D Human Pose and Shape from a Single Color Image
- Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation
- Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision
- Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose
- 3D Human Pose Estimation from a Single Image via Distance Matrix Regression
- Semantic keypoint-based pose estimation from single RGB frames
- DeepSkeleton: Skeleton Map for 3D Human Pose Regression
- Rethinking Pose in 3D: Multi-stage Refinement and Recovery for Markerless Motion Capture
- Ordinal Depth Supervision for 3D Human Pose Estimation
- Deep Non-Rigid Structure from Motion with Missing Data
- Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB
- Recurrent 3D Pose Sequence Machines