Appearance-and-Relation Networks for Video Classification
arXiv:1711.09125
Abstract
Spatiotemporal feature learning in videos is a fundamental problem in computer vision. This paper presents a new architecture, termed as Appearance-and-Relation Network (ARTNet), to learn video representation in an end-to-end manner. ARTNets are constructed by stacking multiple generic building blocks, called as SMART, whose goal is to simultaneously model appearance and relation from RGB input in a separate and explicit manner. Specifically, SMART blocks decouple the spatiotemporal learning module into an appearance branch for spatial modeling and a relation branch for temporal modeling. The appearance branch is implemented based on the linear combination of pixels or filter responses in each frame, while the relation branch is designed based on the multiplicative interactions between pixels or filter responses across multiple frames. We perform experiments on three action recognition benchmarks: Kinetics, UCF101, and HMDB51, demonstrating that SMART blocks obtain an evident improvement over 3D convolutions for spatiotemporal feature learning. Under the same training setting, ARTNets achieve superior performance on these three datasets to the existing state-of-the-art methods.
CVPR18 camera-ready version. Code & models available at https://github.com/wanglimin/ARTNet
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- The Kinetics Human Action Video Dataset
- Going Deeper with Convolutions
- Spatiotemporal Residual Networks for Video Action Recognition
- ConvNet Architecture Search for Spatiotemporal Feature Learning
- Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
- ActivityNet Challenge 2017 Summary
- On multi-view feature learning
Cited by in corpus (9)
- TSM: Temporal Shift Module for Efficient Video Understanding
- StNet: Local and Global Spatial-Temporal Modeling for Action Recognition
- Exploiting Spatial-Temporal Modelling and Multi-Modal Fusion for Human Action Recognition
- Multi-Fiber Networks for Video Recognition
- Learning Representative Temporal Features for Action Recognition
- Motion Feature Network: Fixed Motion Filter for Action Recognition
- Dynamic Graph Modules for Modeling Object-Object Interactions in Activity Recognition
- Human Activity Recognition for Edge Devices
- Discriminative Video Representation Learning Using Support Vector Classifiers