Context-Aware Deep Spatio-Temporal Network for Hand Pose Estimation from Depth Images
arXiv:1810.02994 · doi:10.1109/TCYB.2018.2873733
Abstract
As a fundamental and challenging problem in computer vision, hand pose estimation aims to estimate the hand joint locations from depth images. Typically, the problem is modeled as learning a mapping function from images to hand joint coordinates in a data-driven manner. In this paper, we propose Context-Aware Deep Spatio-Temporal Network (CADSTN), a novel method to jointly model the spatio-temporal properties for hand pose estimation. Our proposed network is able to learn the representations of the spatial information and the temporal structure from the image sequences. Moreover, by adopting adaptive fusion method, the model is capable of dynamically weighting different predictions to lay emphasis on sufficient context. Our method is examined on two common benchmarks, the experimental results demonstrate that our proposed approach achieves the best or the second-best performance with state-of-the-art methods and runs in 60fps.
IEEE Transactions On Cybernetics
References in corpus (13)
- Adam: A Method for Stochastic Optimization
- Auto-Encoding Variational Bayes
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Hands Deep in Deep Learning for Hand Pose Estimation
- Learning to Execute
- Real-time Hand Tracking under Occlusion from an Egocentric RGB-D Sensor
- Pose Guided Structured Region Ensemble Network for Cascaded Hand Pose Estimation
- Region Ensemble Network: Improving Convolutional Network for Hand Pose Estimation
- Hand3D: Hand Pose Estimation using 3D Neural Network
- Articulated Hand Pose Estimation Review
- Direction matters: hand pose estimation from local surface normals
- Learning to Search on Manifolds for 3D Pose Estimation of Articulated Objects