ActionVLAD: Learning spatio-temporal aggregation for action classification
arXiv:1704.02895
Abstract
In this work, we introduce a new video representation for action classification that aggregates local convolutional features across the entire spatio-temporal extent of the video. We do so by integrating state-of-the-art two-stream networks with learnable spatio-temporal feature aggregation. The resulting architecture is end-to-end trainable for whole-video classification. We investigate different strategies for pooling across space and time and combining signals from the different streams. We find that: (i) it is important to pool jointly across space and time, but (ii) appearance and motion streams are best aggregated into their own separate representations. Finally, we show that our representation outperforms the two-stream base architecture by a large margin (13% relative) as well as out-performs other baselines with comparable base architectures on HMDB51, UCF101, and Charades video classification benchmarks.
Accepted to CVPR 2017. Project page: https://rohitgirdhar.github.io/ActionVLAD/
References in corpus (9)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Spatiotemporal Residual Networks for Video Action Recognition
- Towards Good Practices for Very Deep Two-Stream ConvNets
- Convolutional Two-Stream Network Fusion for Video Action Recognition
- Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding
- Long-term Temporal Convolutions for Action Recognition
- A Discriminative CNN Video Representation for Event Detection