Recovering 3D Human Mesh from Monocular Images: A Survey
arXiv:2203.01923 · doi:10.1109/TPAMI.2023.3298850
Abstract
Estimating human pose and shape from monocular images is a long-standing problem in computer vision. Since the release of statistical body models, 3D human mesh recovery has been drawing broader attention. With the same goal of obtaining well-aligned and physically plausible mesh results, two paradigms have been developed to overcome challenges in the 2D-to-3D lifting process: i) an optimization-based paradigm, where different data terms and regularization terms are exploited as optimization objectives; and ii) a regression-based paradigm, where deep learning techniques are embraced to solve the problem in an end-to-end fashion. Meanwhile, continuous efforts are devoted to improving the quality of 3D mesh labels for a wide range of datasets. Though remarkable progress has been achieved in the past decade, the task is still challenging due to flexible body motions, diverse appearances, complex environments, and insufficient in-the-wild annotations. To the best of our knowledge, this is the first survey that focuses on the task of monocular 3D human mesh recovery. We start with the introduction of body models and then elaborate recovery frameworks and training objectives by providing in-depth analyses of their strengths and weaknesses. We also summarize datasets, evaluation metrics, and benchmark results. Open issues and future directions are discussed in the end, hoping to motivate researchers and facilitate their research in this area. A regularly updated project page can be found at https://github.com/tinatiansjz/hmr-survey.
Published in IEEE TPAMI, Survey on monocular 3D human mesh recovery, Project page: https://github.com/tinatiansjz/hmr-survey
References in corpus (7)
- Embodied Hands: Modeling and Capturing Hands and Bodies Together
- Monocular Human Pose Estimation: A Survey of Deep Learning-based Methods
- STAR: Sparse Trained Articulated Human Body Regressor
- PyMAF-X: Towards Well-aligned Full-body Model Regression from Monocular Images
- DenseBody: Directly Regressing Dense 3D Human Pose and Shape From a Single Color Image
- HybrIK-X: Hybrid Analytical-Neural Inverse Kinematics for Whole-body Mesh Recovery
- KBody: Towards general, robust, and aligned monocular whole-body estimation
Cited by in corpus (8)
- PyMAF-X: Towards Well-aligned Full-body Model Regression from Monocular Images
- Deep learning for 3D human pose estimation and mesh recovery: A survey
- Delving Deep into Pixel Alignment Feature for Accurate Multi-view Human Mesh Recovery
- KBody: Towards general, robust, and aligned monocular whole-body estimation
- ARTS: Semi-Analytical Regressor using Disentangled Skeletal Representations for Human Mesh Recovery from Videos
- BundleMoCap: Efficient, Robust and Smooth Motion Capture from Sparse Multiview Videos
- ClothHMR: 3D Mesh Recovery of Humans in Diverse Clothing from Single Image
- Shape-from-Template with Generalised Camera