Human Action Recognition using Factorized Spatio-Temporal Convolutional Networks
arXiv:1510.00562
Abstract
Human actions in video sequences are three-dimensional (3D) spatio-temporal signals characterizing both the visual appearance and motion dynamics of the involved humans and objects. Inspired by the success of convolutional neural networks (CNN) for image classification, recent attempts have been made to learn 3D CNNs for recognizing human actions in videos. However, partly due to the high complexity of training 3D convolution kernels and the need for large quantities of training videos, only limited success has been reported. This has triggered us to investigate in this paper a new deep architecture which can handle 3D signals more effectively. Specifically, we propose factorized spatio-temporal convolutional networks (FstCN) that factorize the original 3D convolution kernel learning as a sequential process of learning 2D spatial kernels in the lower layers (called spatial convolutional layers), followed by learning 1D temporal kernels in the upper layers (called temporal convolutional layers). We introduce a novel transformation and permutation operator to make factorization in FstCN possible. Moreover, to address the issue of sequence alignment, we propose an effective training and inference strategy based on sampling multiple video clips from a given action video sequence. We have tested FstCN on two commonly used benchmark datasets (UCF-101 and HMDB-51). Without using auxiliary training videos to boost the performance, FstCN outperforms existing CNN based methods and achieves comparable performance with a recent method that benefits from using auxiliary training videos.
References in corpus (5)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Going Deeper with Convolutions
- Bag of Visual Words and Fusion Methods for Action Recognition: Comprehensive Study and Good Practice
Cited by in corpus (33)
- Convolutional Two-Stream Network Fusion for Video Action Recognition
- Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
- Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
- Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification
- Flow-Guided Feature Aggregation for Video Object Detection
- Applying Deep Bidirectional LSTM and Mixture Density Network for Basketball Trajectory Prediction
- Tube Convolutional Neural Network (T-CNN) for Action Detection in Videos
- Appearance-and-Relation Networks for Video Classification
- Impression Network for Video Object Detection
- Optical Flow Guided Feature: A Fast and Robust Motion Representation for Video Action Recognition
- V2CNet: A Deep Learning Framework to Translate Videos to Commands for Robotic Manipulation
- Temporal Convolutional Networks for Action Segmentation and Detection
- Lattice Long Short-Term Memory for Human Action Recognition
- Attention Clusters: Purely Attention Based Local Feature Integration for Video Classification
- Attention-based Temporal Weighted Convolutional Neural Network for Action Recognition
- Semantic Video Segmentation by Gated Recurrent Flow Propagation
- End-to-end Video-level Representation Learning for Action Recognition
- Deep Temporal Linear Encoding Networks
- Asynchronous Temporal Fields for Action Recognition
- Learning to score the figure skating sports videos
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- Two-stream Collaborative Learning with Spatial-Temporal Attention for Video Classification
- Self-Supervised Video Representation Learning With Odd-One-Out Networks
- Learning Gating ConvNet for Two-Stream based Methods in Action Recognition
- Learning long-term dependencies for action recognition with a biologically-inspired deep network
- ActionVLAD: Learning spatio-temporal aggregation for action classification
- Discriminatively Learned Hierarchical Rank Pooling Networks
- Improved Dense Trajectory with Cross Streams
- Temporal Bilinear Networks for Video Action Recognition
- Temporal-Spatial Mapping for Action Recognition
- Deep Discriminative Model for Video Classification
- Temporal Factorization of 3D Convolutional Kernels
- Segmental Spatiotemporal CNNs for Fine-grained Action Segmentation