Vision-Based Human Pose Estimation via Deep Learning: A Survey
arXiv:2308.13872 · doi:10.1109/THMS.2022.3219242
Abstract
Human pose estimation (HPE) has attracted a significant amount of attention from the computer vision community in the past decades. Moreover, HPE has been applied to various domains, such as human-computer interaction, sports analysis, and human tracking via images and videos. Recently, deep learning-based approaches have shown state-of-the-art performance in HPE-based applications. Although deep learning-based approaches have achieved remarkable performance in HPE, a comprehensive review of deep learning-based HPE methods remains lacking in the literature. In this article, we provide an up-to-date and in-depth overview of the deep learning approaches in vision-based HPE. We summarize these methods of 2-D and 3-D HPE, and their applications, discuss the challenges and the research trends through bibliometrics, and provide insightful recommendations for future research. This article provides a meaningful overview as introductory material for beginners to deep learning-based HPE, as well as supplementary material for advanced researchers.
16 pages, 4 figures
References in corpus (6)
- NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding
- Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation
- Monocular Human Pose Estimation: A Survey of Deep Learning-based Methods
- AdaFuse: Adaptive Multiview Fusion for Accurate Human Pose Estimation in the Wild
- TFPose: Direct Human Pose Estimation with Transformers
- Recent Advances in Monocular 2D and 3D Human Pose Estimation: A Deep Learning Perspective