An End-to-End Spatio-Temporal Attention Model for Human Action Recognition from Skeleton Data
arXiv:1611.06067
Abstract
Human action recognition is an important task in computer vision. Extracting discriminative spatial and temporal features to model the spatial and temporal evolutions of different actions plays a key role in accomplishing this task. In this work, we propose an end-to-end spatial and temporal attention model for human action recognition from skeleton data. We build our model on top of the Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM), which learns to selectively focus on discriminative joints of skeleton within each frame of the inputs and pays different levels of attention to the outputs of different frames. Furthermore, to ensure effective training of the network, we propose a regularized cross-entropy loss to drive the model learning process and develop a joint training strategy accordingly. Experimental results demonstrate the effectiveness of the proposed model,both on the small human action recognition data set of SBU and the currently largest NTU dataset.
References in corpus (2)
Cited by in corpus (41)
- Attentional Pooling for Action Recognition
- Multi-task Deep Learning for Real-Time 3D Human Pose Estimation and Action Recognition
- An Attention Enhanced Graph Convolutional LSTM Network for Skeleton-Based Action Recognition
- Empowering Things with Intelligence: A Survey of the Progress, Challenges, and Opportunities in Artificial Intelligence of Things
- Learning Graph Convolutional Network for Skeleton-based Human Action Recognition by Neural Searching
- Complex Sequential Understanding through the Awareness of Spatial and Temporal Concepts
- UNIK: A Unified Framework for Real-world Skeleton-based Action Recognition
- Exploiting the ConvLSTM: Human Action Recognition using Raw Depth Video-Based Recurrent Neural Networks
- Detailed 2D-3D Joint Representation for Human-Object Interaction
- MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion Prediction
- Focusing and Diffusion: Bidirectional Attentive Graph Convolutional Networks for Skeleton-based Action Recognition
- Video Super-resolution with Temporal Group Attention
- ElderSim: A Synthetic Data Generation Platform for Human Action Recognition in Eldercare Applications
- CTCModel: a Keras Model for Connectionist Temporal Classification
- A Two-stream Neural Network for Pose-based Hand Gesture Recognition
- Understanding the Robustness of Skeleton-based Action Recognition under Adversarial Attack
- HAN: An Efficient Hierarchical Self-Attention Network for Skeleton-Based Gesture Recognition
- Group-Skeleton-Based Human Action Recognition in Complex Events
- Skeleton based Activity Recognition by Fusing Part-wise Spatio-temporal and Attention Driven Residues
- Towards Coding for Human and Machine Vision: A Scalable Image Coding Approach
- Action Recognition with Coarse-to-Fine Deep Feature Integration and Asynchronous Fusion
- Mix Dimension in Poincaré Geometry for 3D Skeleton-based Action Recognition
- Human Action Recognition with Multi-Laplacian Graph Convolutional Networks
- Cross-modal knowledge distillation for action recognition
- Multi-Scale Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition
- Selective Spatio-Temporal Aggregation Based Pose Refinement System: Towards Understanding Human Activities in Real-World Videos
- iMiGUE: An Identity-free Video Dataset for Micro-Gesture Understanding and Emotion Analysis
- An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated Videos
- Multi Scale Temporal Graph Networks For Skeleton-based Action Recognition
- Learning Various Length Dependence by Dual Recurrent Neural Networks
- Learning Coupled Spatial-temporal Attention for Skeleton-based Action Recognition
- Coarse Temporal Attention Network (CTA-Net) for Driver's Activity Recognition
- Improving Skeleton-based Action Recognitionwith Robust Spatial and Temporal Features
- Totally Deep Support Vector Machines
- SAR-NAS: Skeleton-based Action Recognition via Neural Architecture Searching
- An Emerging Coding Paradigm VCM: A Scalable Coding Approach Beyond Feature and Signal
- Action Recognition with Kernel-based Graph Convolutional Networks
- Modality Compensation Network: Cross-Modal Adaptation for Action Recognition
- Video-Based Convolutional Attention for Person Re-Identification
- Attention-Driven Body Pose Encoding for Human Activity Recognition
- Object Properties Inferring from and Transfer for Human Interaction Motions