End-to-end Global to Local CNN Learning for Hand Pose Recovery in Depth Data
arXiv:1705.09606
Abstract
Despite recent advances in 3D pose estimation of human hands, especially thanks to the advent of CNNs and depth cameras, this task is still far from being solved. This is mainly due to the highly non-linear dynamics of fingers, which make hand model training a challenging task. In this paper, we exploit a novel hierarchical tree-like structured CNN, in which branches are trained to become specialized in predefined subsets of hand joints, called local poses. We further fuse local pose features, extracted from hierarchical CNN branches, to learn higher order dependencies among joints in the final pose by end-to-end training. Lastly, the loss function used is also defined to incorporate appearance and physical constraints about doable hand motion and deformation. Finally, we introduce a non-rigid data augmentation approach to increase the amount of training depth data. Experimental results suggest that feeding a tree-shaped CNN, specialized in local poses, into a fusion network for modeling joints correlations and dependencies, helps to increase the precision of final estimations, outperforming state-of-the-art results on NYU and SyntheticHand datasets.
References in corpus (2)
Cited by in corpus (11)
- Pose Guided Structured Region Ensemble Network for Cascaded Hand Pose Estimation
- V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation from a Single Depth Map
- A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation from a Single Depth Image
- Fast and Accurate 3D Hand Pose Estimation via Recurrent Neural Network for Capturing Hand Articulations
- HandAugment: A Simple Data Augmentation Method for Depth-Based 3D Hand Pose Estimation
- Simultaneous Hand Pose and Skeleton Bone-Lengths Estimation from a Single Depth Image
- Model-based Hand Pose Estimation for Generalized Hand Shape with Appearance Normalization
- Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals
- JGR-P2O: Joint Graph Reasoning based Pixel-to-Offset Prediction Network for 3D Hand Pose Estimation from a Single Depth Image
- Silhouette-Net: 3D Hand Pose Estimation from Silhouettes
- HMTNet:3D Hand Pose Estimation from Single Depth Image Based on Hand Morphological Topology