3D Human Pose Estimation = 2D Pose Estimation + Matching
arXiv:1612.06524
Abstract
We explore 3D human pose estimation from a single RGB image. While many approaches try to directly predict 3D pose from image measurements, we explore a simple architecture that reasons through intermediate 2D pose predictions. Our approach is based on two key observations (1) Deep neural nets have revolutionized 2D pose estimation, producing accurate 2D predictions even for poses with self occlusions. (2) Big-data sets of 3D mocap data are now readily available, making it tempting to lift predicted 2D poses to 3D through simple memorization (e.g., nearest neighbors). The resulting architecture is trivial to implement with off-the-shelf 2D pose estimation systems and 3D mocap libraries. Importantly, we demonstrate that such methods outperform almost all state-of-the-art 3D pose estimation systems, most of which directly try to regress 3D pose from 2D measurements.
Demo code: https://github.com/flyawaychase/3DHumanPose
References in corpus (2)
Cited by in corpus (9)
- Self-supervised Learning of Motion Capture
- Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach
- Passing a Non-verbal Turing Test: Evaluating Gesture Animations Generated from Speech
- Learning Pose Grammar to Encode Human Body Configuration for 3D Pose Estimation
- Integral Human Pose Regression
- DeepSkeleton: Skeleton Map for 3D Human Pose Regression
- Toward Marker-free 3D Pose Estimation in Lifting: A Deep Multi-view Solution
- PedX: Benchmark Dataset for Metric 3D Pose Estimation of Pedestrians in Complex Urban Intersections
- Explicit Spatiotemporal Joint Relation Learning for Tracking Human Pose