Unite the People: Closing the Loop Between 3D and 2D Human Representations
arXiv:1701.02468
Abstract
3D models provide a common ground for different representations of human bodies. In turn, robust 2D estimation has proven to be a powerful tool to obtain 3D fits "in-the- wild". However, depending on the level of detail, it can be hard to impossible to acquire labeled data for training 2D estimators on large scale. We propose a hybrid approach to this problem: with an extended version of the recently introduced SMPLify method, we obtain high quality 3D body model fits for multiple human pose datasets. Human annotators solely sort good and bad fits. This procedure leads to an initial dataset, UP-3D, with rich annotations. With a comprehensive set of experiments, we show how this data can be used to train discriminative models that produce results with an unprecedented level of detail: our models predict 31 segments and 91 landmark locations on the body. Using the 91 landmark pose estimator, we present state-of-the art results for 3D human pose and shape estimation using an order of magnitude less training data and without assumptions about gender or pose in the fitting procedure. We show that UP-3D can be enhanced with these improved fits to grow in quantity and quality, which makes the system deployable on large scale. The data, code and models are available for research purposes.
References in corpus (1)
Cited by in corpus (20)
- Efficient Convolutional Neural Networks for Depth-Based Multi-Person Pose Estimation
- 3D Hand Shape and Pose Estimation from a Single RGB Image
- 3D Human Pose Estimation with Relational Networks
- Video Based Reconstruction of 3D People Models
- Detailed, accurate, human shape estimation from clothed 3D scan sequences
- Learning to Estimate 3D Human Pose and Shape from a Single Color Image
- A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and Pose
- Single Image 3D Hand Reconstruction with Mesh Convolutions
- Towards Robust RGB-D Human Mesh Recovery
- Learning 3D Human Dynamics from Video
- A Neural Anthropometer Learning from Body Dimensions Computed on Human 3D Meshes
- Adaloss: Adaptive Loss Function for Landmark Localization
- Canonical 3D Deformer Maps: Unifying parametric and non-parametric methods for dense weakly-supervised category reconstruction
- Ordinal Depth Supervision for 3D Human Pose Estimation
- ChallenCap: Monocular 3D Capture of Challenging Human Performances using Multi-Modal References
- Fashion is Taking Shape: Understanding Clothing Preference Based on Body Shape From Online Sources
- LBS Autoencoder: Self-supervised Fitting of Articulated Meshes to Point Clouds
- Coherent Reconstruction of Multiple Humans from a Single Image
- Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB
- PedX: Benchmark Dataset for Metric 3D Pose Estimation of Pedestrians in Complex Urban Intersections