Skeleton based action recognition using translation-scale invariant image mapping and multi-scale deep cnn
arXiv:1704.05645
Abstract
This paper presents an image classification based approach for skeleton-based video action recognition problem. Firstly, A dataset independent translation-scale invariant image mapping method is proposed, which transformes the skeleton videos to colour images, named skeleton-images. Secondly, A multi-scale deep convolutional neural network (CNN) architecture is proposed which could be built and fine-tuned on the powerful pre-trained CNNs, e.g., AlexNet, VGGNet, ResNet etal.. Even though the skeleton-images are very different from natural images, the fine-tune strategy still works well. At last, we prove that our method could also work well on 2D skeleton video data. We achieve the state-of-the-art results on the popular benchmard datasets e.g. NTU RGB+D, UTD-MHAD, MSRC-12, and G3D. Especially on the largest and challenge NTU RGB+D, UTD-MHAD, and MSRC-12 dataset, our method outperforms other methods by a large margion, which proves the efficacy of the proposed method.
References in corpus (3)
Cited by in corpus (14)
- Monocular Human Pose Estimation: A Survey of Deep Learning-based Methods
- Skeleton-based Action Recognition via Spatial and Temporal Transformer Networks
- Richly Activated Graph Convolutional Network for Robust Skeleton-based Action Recognition
- Infrared and 3D skeleton feature fusion for RGB-D action recognition
- View Adaptive Neural Networks for High Performance Skeleton-based Human Action Recognition
- Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition
- Feedback Graph Convolutional Network for Skeleton-based Action Recognition
- Predictively Encoded Graph Convolutional Network for Noise-Robust Skeleton-based Action Recognition
- Temporal Extension Module for Skeleton-Based Action Recognition
- Improving Skeleton-based Action Recognitionwith Robust Spatial and Temporal Features
- Spatio-Temporal Dual Affine Differential Invariant for Skeleton-based Action Recognition
- Attention-Driven Body Pose Encoding for Human Activity Recognition
- Skeleton-based Approaches based on Machine Vision: A Survey
- DeepActsNet: Spatial and Motion features from Face, Hands, and Body Combined with Convolutional and Graph Networks for Improved Action Recognition