A simple yet effective baseline for 3d human pose estimation
arXiv:1705.03098
Abstract
Following the success of deep convolutional networks, state-of-the-art methods for 3d human pose estimation have focused on deep end-to-end systems that predict 3d joint locations given raw image pixels. Despite their excellent performance, it is often not easy to understand whether their remaining error stems from a limited 2d pose (visual) understanding, or from a failure to map 2d poses into 3-dimensional positions. With the goal of understanding these sources of error, we set out to build a system that given 2d joint locations predicts 3d positions. Much to our surprise, we have found that, with current technology, "lifting" ground truth 2d joint locations to 3d space is a task that can be solved with a remarkably low error rate: a relatively simple deep feed-forward network outperforms the best reported result by about 30\% on Human3.6M, the largest publicly available 3d pose estimation benchmark. Furthermore, training our system on the output of an off-the-shelf state-of-the-art 2d detector (\ie, using images as input) yields state of the art results -- this includes an array of systems that have been trained end-to-end specifically for this task. Our results indicate that a large portion of the error of modern deep 3d pose estimation systems stems from their visual analysis, and suggests directions to further advance the state of the art in 3d human pose estimation.
Accepted to ICCV 17
References in corpus (2)
Cited by in corpus (33)
- Weakly-Supervised Discovery of Geometry-Aware Representation for 3D Human Pose Estimation
- Real-Time Human Pose Estimation on a Smart Walker using Convolutional Neural Networks
- Fast and Robust Multi-Person 3D Pose Estimation from Multiple Views
- Learning Pose Grammar to Encode Human Body Configuration for 3D Pose Estimation
- Monocular Total Capture: Posing Face, Body, and Hands in the Wild
- Learning to Estimate 3D Human Pose and Shape from a Single Color Image
- Inference Stage Optimization for Cross-scenario 3D Human Pose Estimation
- Integral Human Pose Regression
- PoseAug: A Differentiable Pose Augmentation Framework for 3D Human Pose Estimation
- Multi-view Human Pose and Shape Estimation Using Learnable Volumetric Aggregation
- FBI-Pose: Towards Bridging the Gap between 2D Images and 3D Human Poses using Forward-or-Backward Information
- Towards Robust RGB-D Human Mesh Recovery
- DRPose3D: Depth Ranking in 3D Human Pose Estimation
- Cross View Fusion for 3D Human Pose Estimation
- What Face and Body Shapes Can Tell About Height
- 3D Human Pose Estimation with 2D Marginal Heatmaps
- DeepSkeleton: Skeleton Map for 3D Human Pose Regression
- Recovering and Simulating Pedestrians in the Wild
- On the Robustness of Human Pose Estimation
- Rethinking Pose in 3D: Multi-stage Refinement and Recovery for Markerless Motion Capture
- Higher-Order Implicit Fairing Networks for 3D Human Pose Estimation
- Dense 3D Regression for Hand Pose Estimation
- Ordinal Depth Supervision for 3D Human Pose Estimation
- Post-Data Augmentation to Improve Deep Pose Estimation of Extreme and Wild Motions
- PI-Net: Pose Interacting Network for Multi-Person Monocular 3D Pose Estimation
- Distill Knowledge from NRSfM for Weakly Supervised 3D Pose Learning
- View Invariant 3D Human Pose Estimation
- Structure from Recurrent Motion: From Rigidity to Recurrency
- PedX: Benchmark Dataset for Metric 3D Pose Estimation of Pedestrians in Complex Urban Intersections
- Learning to Predict Diverse Human Motions from a Single Image via Mixture Density Networks
- Coherent Reconstruction of Multiple Humans from a Single Image
- Can Human Sex Be Learned Using Only 2D Body Keypoint Estimations?
- Explicit Spatiotemporal Joint Relation Learning for Tracking Human Pose