End-to-end Video-level Representation Learning for Action Recognition
arXiv:1711.04161
Abstract
From the frame/clip-level feature learning to the video-level representation building, deep learning methods in action recognition have developed rapidly in recent years. However, current methods suffer from the confusion caused by partial observation training, or without end-to-end learning, or restricted to single temporal scale modeling and so on. In this paper, we build upon two-stream ConvNets and propose Deep networks with Temporal Pyramid Pooling (DTPP), an end-to-end video-level representation learning approach, to address these problems. Specifically, at first, RGB images and optical flow stacks are sparsely sampled across the whole video. Then a temporal pyramid pooling layer is used to aggregate the frame-level features which consist of spatial and temporal cues. Lastly, the trained model has compact video-level representation with multiple temporal scales, which is both global and sequence-aware. Experimental results show that DTPP achieves the state-of-the-art performance on two challenging video action datasets: UCF101 and HMDB51, either by ImageNet pre-training or Kinetics pre-training.
10 pages, 6 figures, 6 tables. The explanation for the batch size is added. Accepted by ICPR 2018
References in corpus (16)
- Sequence to Sequence Learning with Neural Networks
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Spatiotemporal Residual Networks for Video Action Recognition
- Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
- Show and Tell: A Neural Image Caption Generator
- Temporal Segment Networks for Action Recognition in Videos
- FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks
- TS-LSTM and Temporal-Inception: Exploiting Spatiotemporal Dynamics for Activity Recognition
- Temporal Pyramid Pooling Based Convolutional Neural Networks for Action Recognition
- Beyond Gaussian Pyramid: Multi-skip Feature Stacking for Action Recognition
- Action Recognition with Dynamic Image Networks
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- Eigen Evolution Pooling for Human Action Recognition
- Learning Gating ConvNet for Two-Stream based Methods in Action Recognition