Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision
arXiv:1611.09813
Abstract
We propose a CNN-based approach for 3D human body pose estimation from single RGB images that addresses the issue of limited generalizability of models trained solely on the starkly limited publicly available 3D pose data. Using only the existing 3D pose data and 2D pose data, we show state-of-the-art performance on established benchmarks through transfer of learned features, while also generalizing to in-the-wild scenes. We further introduce a new training set for human body pose estimation from monocular images of real humans that has the ground truth captured with a multi-camera marker-less motion capture system. It complements existing corpora with greater diversity in pose, human appearance, clothing, occlusion, and viewpoints, and enables an increased scope of augmentation. We also contribute a new benchmark that covers outdoor and indoor scenes, and demonstrate that our 3D pose dataset shows better in-the-wild performance than existing annotated data, which is further improved in conjunction with transfer learning from 2D pose data. All in all, we argue that the use of transfer learning of representations in tandem with algorithmic and data contributions is crucial for general 3D body pose estimation.
Accepted at the International Conference on 3D Vision (3DV) 2017
References in corpus (17)
- Deep Residual Learning for Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Unsupervised Domain Adaptation by Backpropagation
- Fully Convolutional Networks for Semantic Segmentation
- Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image
- Human pose estimation via Convolutional Part Heatmap Regression
- Stacked Hourglass Networks for Human Pose Estimation
- MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild
- Synthesizing Training Images for Boosting Human 3D Pose Estimation
- Structured Prediction of 3D Human Pose with Deep Neural Networks
- MoDeep: A Deep Learning Framework Using Motion Features for Human Pose Estimation
- Human Pose Estimation using Deep Consensus Voting
- General Automatic Human Shape and Motion Capture Using Volumetric Contour Cues
- Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation
- 3D Human Pose Estimation from a Single Image via Distance Matrix Regression
- 3D Shape Estimation from 2D Landmarks: A Convex Relaxation Approach
- Model-based Outdoor Performance Capture
Cited by in corpus (14)
- VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera
- Real-time Hand Tracking under Occlusion from an Egocentric RGB-D Sensor
- Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach
- MonoPerfCap: Human Performance Capture from Monocular Video
- 3D Human Pose Estimation in the Wild by Adversarial Learning
- Learning to Estimate 3D Human Pose and Shape from a Single Color Image
- Integral Human Pose Regression
- DRPose3D: Depth Ranking in 3D Human Pose Estimation
- DeepSkeleton: Skeleton Map for 3D Human Pose Regression
- Ordinal Depth Supervision for 3D Human Pose Estimation
- PedX: Benchmark Dataset for Metric 3D Pose Estimation of Pedestrians in Complex Urban Intersections
- Skeleton Transformer Networks: 3D Human Pose and Skinned Mesh from Single RGB Image
- Explicit Spatiotemporal Joint Relation Learning for Tracking Human Pose
- MEBOW: Monocular Estimation of Body Orientation In the Wild