RGB-D-based Human Motion Recognition with Deep Learning: A Survey
arXiv:1711.08362
Abstract
Human motion recognition is one of the most important branches of human-centered research activities. In recent years, motion recognition based on RGB-D data has attracted much attention. Along with the development in artificial intelligence, deep learning techniques have gained remarkable success in computer vision. In particular, convolutional neural networks (CNN) have achieved great success for image-based tasks, and recurrent neural networks (RNN) are renowned for sequence-based problems. Specifically, deep learning methods based on the CNN and RNN architectures have been adopted for motion recognition using RGB-D data. In this paper, a detailed overview of recent advances in RGB-D-based motion recognition is presented. The reviewed methods are broadly categorized into four groups, depending on the modality adopted for recognition: RGB-based, depth-based, skeleton-based and RGB+D-based. As a survey focused on the application of deep learning to RGB-D-based motion recognition, we explicitly discuss the advantages and limitations of existing techniques. Particularly, we highlighted the methods of encoding spatial-temporal-structural information inherent in video sequence, and discuss potential directions for future research.
References in corpus (21)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Generating Videos with Scene Dynamics
- Spatiotemporal Residual Networks for Video Action Recognition
- An End-to-End Spatio-Temporal Attention Model for Human Action Recognition from Skeleton Data
- PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding
- Temporal Action Detection with Structured Segment Networks
- Differential Recurrent Neural Networks for Action Recognition
- Joint Geometrical and Statistical Alignment for Visual Domain Adaptation
- TURN TAP: Temporal Unit Regression Network for Temporal Action Proposals
- Tube Convolutional Neural Network (T-CNN) for Action Detection in Videos
- Cross-view Action Modeling, Learning and Recognition
- View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition from Skeleton Data
- Cooperative Training of Deep Aggregation Networks for RGB-D Action Recognition
- Predictive-Corrective Networks for Action Detection
- Multi-Modality Fusion based on Consensus-Voting and 3D Convolution for Isolated Gesture Recognition
- Unsupervised Learning of Long-Term Motion Dynamics for Videos
- Lattice Long Short-Term Memory for Human Action Recognition
- Generalized Rank Pooling for Activity Recognition
- Large-scale Continuous Gesture Recognition Using Convolutional Neural Networks
- AMTnet: Action-Micro-Tube Regression by End-to-end Trainable Deep Architecture