Skeleton-based Action Recognition of People Handling Objects
arXiv:1901.06882 · doi:10.1109/WACV.2019.00014
Abstract
In visual surveillance systems, it is necessary to recognize the behavior of people handling objects such as a phone, a cup, or a plastic bag. In this paper, to address this problem, we propose a new framework for recognizing object-related human actions by graph convolutional networks using human and object poses. In this framework, we construct skeletal graphs of reliable human poses by selectively sampling the informative frames in a video, which include human joints with high confidence scores obtained in pose estimation. The skeletal graphs generated from the sampled frames represent human poses related to the object position in both the spatial and temporal domains, and these graphs are used as inputs to the graph convolutional networks. Through experiments over an open benchmark and our own data sets, we verify the validity of our framework in that our method outperforms the state-of-the-art method for skeleton-based action recognition.
Accepted in WACV 2019
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- A New Representation of Skeleton Sequences for 3D Action Recognition
- NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis
- Convolutional Two-Stream Network Fusion for Video Action Recognition
- Temporal Segment Networks: Towards Good Practices for Deep Action Recognition