Video action detection by learning graph-based spatio-temporal interactions
arXiv:1912.04316 · doi:10.1016/j.cviu.2021.103187
Abstract
Action Detection is a complex task that aims to detect and classify human actions in video clips. Typically, it has been addressed by processing fine-grained features extracted from a video classification backbone. Recently, thanks to the robustness of object and people detectors, a deeper focus has been added on relationship modelling. Following this line, we propose a graph-based framework to learn high-level interactions between people and objects, in both space and time. In our formulation, spatio-temporal relationships are learned through self-attention on a multi-layer graph structure which can connect entities from consecutive clips, thus considering long-range spatial and temporal dependencies. The proposed module is backbone independent by design and does not require end-to-end training. Extensive experiments are conducted on the AVA dataset, where our model demonstrates state-of-the-art results and consistent improvements over baselines built with different backbones. Code is publicly available at https://github.com/aimagelab/STAGE_action_detection.
This is the authors version of an article accepted for publication in Computer Vision and Image Understanding (CVIU), available online February 2021
References in corpus (12)
- Adam: A Method for Stochastic Optimization
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Semi-Supervised Classification with Graph Convolutional Networks
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- The Kinetics Human Action Video Dataset
- Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Spatiotemporal Residual Networks for Video Action Recognition
- Long-term Temporal Convolutions for Action Recognition
- A Better Baseline for AVA
- Collaborative Spatio-temporal Feature Learning for Video Action Recognition