Pose-conditioned Spatio-Temporal Attention for Human Action Recognition
arXiv:1703.10106
Abstract
We address human action recognition from multi-modal video data involving articulated pose and RGB frames and propose a two-stream approach. The pose stream is processed with a convolutional model taking as input a 3D tensor holding data from a sub-sequence. A specific joint ordering, which respects the topology of the human body, ensures that different convolutional layers correspond to meaningful levels of abstraction. The raw RGB stream is handled by a spatio-temporal soft-attention mechanism conditioned on features from the pose network. An LSTM network receives input from a set of image locations at each instant. A trainable glimpse sensor extracts features on a set of predefined locations specified by the pose stream, namely the 4 hands of the two people involved in the activity. Appearance features give important cues on hand motion and on objects held in each hand. We show that it is of high interest to shift the attention to different hands at different time steps depending on the activity itself. Finally a temporal attention mechanism learns how to fuse LSTM features over time. We evaluate the method on 3 datasets. State-of-the-art results are achieved on the largest dataset for human activity recognition, namely NTU-RGB+D, as well as on the SBU Kinect Interaction dataset. Performance close to state-of-the-art is achieved on the smaller MSR Daily Activity 3D dataset.
10 pages, project page: https://fabienbaradel.github.io/pose_rgb_attention_human_action
References in corpus (8)
- Recurrent Models of Visual Attention
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Multiple Object Recognition with Visual Attention
- An End-to-End Spatio-Temporal Attention Model for Human Action Recognition from Skeleton Data
- Action Recognition using Visual Attention
- Recurrent Mixture Density Network for Spatiotemporal Visual Attention
- A Deep Structured Model with Radius-Margin Bound for 3D Human Activity Recognition
- Action Recognition Based on Joint Trajectory Maps with Convolutional Neural Networks
Cited by in corpus (13)
- NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding
- Human Action Recognition from Various Data Modalities: A Review
- Multi-task Deep Learning for Real-Time 3D Human Pose Estimation and Action Recognition
- Synthetic Humans for Action Recognition from Unseen Viewpoints
- Infrared and 3D skeleton feature fusion for RGB-D action recognition
- Deep Independently Recurrent Neural Network (IndRNN)
- Pose is all you need: The pose only group activity recognition system (POGARS)
- Learning to recognize touch gestures: recurrent vs. convolutional features and dynamic sampling
- Human Action Recognition with Multi-Laplacian Graph Convolutional Networks
- Totally Deep Support Vector Machines
- DeepActsNet: Spatial and Motion features from Face, Hands, and Body Combined with Convolutional and Graph Networks for Improved Action Recognition
- Action Recognition with Kernel-based Graph Convolutional Networks
- Multi-Modal Three-Stream Network for Action Recognition