Space-Time Representation of People Based on 3D Skeletal Data: A Review
arXiv:1601.01006
Abstract
Spatiotemporal human representation based on 3D visual perception data is a rapidly growing research area. Based on the information sources, these representations can be broadly categorized into two groups based on RGB-D information or 3D skeleton data. Recently, skeleton-based human representations have been intensively studied and kept attracting an increasing attention, due to their robustness to variations of viewpoint, human body scale and motion speed as well as the realtime, online performance. This paper presents a comprehensive survey of existing space-time representations of people based on 3D skeletal data, and provides an informative categorization and analysis of these methods from the perspectives, including information modality, representation encoding, structure and transition, and feature engineering. We also provide a brief overview of skeleton acquisition devices and construction methods, enlist a number of public benchmark datasets with skeleton data, and discuss potential future research directions.
Our paper has been accepted by the journal Computer Vision and Image Understanding, see http://www.sciencedirect.com/science/article/pii/S1077314217300279, Computer Vision and Image Understanding, 2017
References in corpus (4)
- Combining Local Appearance and Holistic View: Dual-Source Deep Neural Networks for Human Pose Estimation
- Cross-view Action Modeling, Learning and Recognition
- A discussion on the validation tests employed to compare human action recognition methods using the MSR Action3D dataset
- Simultaneous Feature and Body-Part Learning for Real-Time Robot Awareness of Human Behaviors
Cited by in corpus (10)
- NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis
- View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition from Skeleton Data
- Deep representation learning for human motion prediction and classification
- Learning Human Motion Models for Long-term Predictions
- Deep Multimodal Feature Analysis for Action Recognition in RGB+D Videos
- Simultaneous Joint and Object Trajectory Templates for Human Activity Recognition from 3-D Data
- Deep Learning on Lie Groups for Skeleton-based Action Recognition
- Multi-Scale Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition
- Spatio-Temporal LSTM with Trust Gates for 3D Human Action Recognition
- Online Human Action Detection using Joint Classification-Regression Recurrent Neural Networks