An End-to-End Spatio-Temporal Attention Model for Human Action Recognition from Skeleton Data
arXiv:1611.06067
Abstract
Human action recognition is an important task in computer vision. Extracting discriminative spatial and temporal features to model the spatial and temporal evolutions of different actions plays a key role in accomplishing this task. In this work, we propose an end-to-end spatial and temporal attention model for human action recognition from skeleton data. We build our model on top of the Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM), which learns to selectively focus on discriminative joints of skeleton within each frame of the inputs and pays different levels of attention to the outputs of different frames. Furthermore, to ensure effective training of the network, we propose a regularized cross-entropy loss to drive the model learning process and develop a joint training strategy accordingly. Experimental results demonstrate the effectiveness of the proposed model,both on the small human action recognition data set of SBU and the currently largest NTU dataset.
References in corpus (2)
Cited by in corpus (17)
- Attentional Pooling for Action Recognition
- Multi-task Deep Learning for Real-Time 3D Human Pose Estimation and Action Recognition
- An Attention Enhanced Graph Convolutional LSTM Network for Skeleton-Based Action Recognition
- Learning Graph Convolutional Network for Skeleton-based Human Action Recognition by Neural Searching
- Focusing and Diffusion: Bidirectional Attentive Graph Convolutional Networks for Skeleton-based Action Recognition
- CTCModel: a Keras Model for Connectionist Temporal Classification
- Towards Coding for Human and Machine Vision: A Scalable Image Coding Approach
- Skeleton based Activity Recognition by Fusing Part-wise Spatio-temporal and Attention Driven Residues
- Human Action Recognition with Multi-Laplacian Graph Convolutional Networks
- Action Recognition with Coarse-to-Fine Deep Feature Integration and Asynchronous Fusion
- Cross-modal knowledge distillation for action recognition
- An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated Videos
- Learning Coupled Spatial-temporal Attention for Skeleton-based Action Recognition
- An Emerging Coding Paradigm VCM: A Scalable Coding Approach Beyond Feature and Signal
- Totally Deep Support Vector Machines
- Video-Based Convolutional Attention for Person Re-Identification
- Modality Compensation Network: Cross-Modal Adaptation for Action Recognition