Deep learning for 3D human pose estimation and mesh recovery: A survey
arXiv:2402.18844 · doi:10.1016/j.neucom.2024.128049
Abstract
3D human pose estimation and mesh recovery have attracted widespread research interest in many areas, such as computer vision, autonomous driving, and robotics. Deep learning on 3D human pose estimation and mesh recovery has recently thrived, with numerous methods proposed to address different problems in this area. In this paper, to stimulate future research, we present a comprehensive review of recent progress over the past five years in deep learning methods for this area by delving into over 200 references. To the best of our knowledge, this survey is arguably the first to comprehensively cover deep learning methods for 3D human pose estimation, including both single-person and multi-person approaches, as well as human mesh recovery, encompassing methods based on explicit models and implicit representations. We also present comparative results on several publicly available datasets, together with insightful observations and inspiring future research directions. A regularly updated project page can be found at https://github.com/liuyangme/SOTA-3DHPE-HMR.
References in corpus (33)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- MLP-Mixer: An all-MLP Architecture for Vision
- Action Transformer: A Self-Attention Model for Short-Time Pose-Based Human Action Recognition
- Skeleton-Aware Networks for Deep Motion Retargeting
- Recovering 3D Human Mesh from Monocular Images: A Survey
- Multi-task Deep Learning for Real-Time 3D Human Pose Estimation and Action Recognition
- PyMAF-X: Towards Well-aligned Full-body Model Regression from Monocular Images
- AdaFuse: Adaptive Multiview Fusion for Accurate Human Pose Estimation in the Wild
- Track Anything: Segment Anything Meets Videos
- Adaptive Multi-view and Temporal Fusing Transformer for 3D Human Pose Estimation
- Regular Splitting Graph Network for 3D Human Pose Estimation
- SparsePoser: Real-time Full-body Motion Reconstruction from Sparse Data
- Learning Dynamical Human-Joint Affinity for 3D Pose Estimation in Videos
- 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
- Dual networks based 3D Multi-Person Pose Estimation from Monocular Video
- SMPLer-X: Scaling Up Expressive Human Pose and Shape Estimation
- LASOR: Learning Accurate 3D Human Pose and Shape Via Synthetic Occlusion-Aware Data and Neural Mesh Rendering
- Temporal Representation Learning on Monocular Videos for 3D Human Pose Estimation
- CaPhy: Capturing Physical Properties for Animatable Human Avatars
- HybrIK-X: Hybrid Analytical-Neural Inverse Kinematics for Whole-body Mesh Recovery
- HSTFormer: Hierarchical Spatial-Temporal Transformers for 3D Human Pose Estimation
- Learning Disentangled Avatars with Hybrid 3D Representations
- GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting
- Deep unsupervised 3D human body reconstruction from a sparse set of landmarks
- Text-to-3D using Gaussian Splatting
- ChatPose: Chatting about 3D Human Pose
- Global-correlated 3D-decoupling Transformer for Clothed Avatar Reconstruction
- Live Stream Temporally Embedded 3D Human Body Pose and Shape Estimation
- HumanGaussian: Text-Driven 3D Human Generation with Gaussian Splatting
- Animatable 3D Gaussian: Fast and High-Quality Reconstruction of Multiple Human Avatars
- Morphable Diffusion: 3D-Consistent Diffusion for Single-image Avatar Creation
- MPS-NeRF: Generalizable 3D Human Rendering from Multiview Images
- HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video