Human Action Localization with Sparse Spatial Supervision
arXiv:1605.05197
Abstract
We introduce an approach for spatio-temporal human action localization using sparse spatial supervision. Our method leverages the large amount of annotated humans available today and extracts human tubes by combining a state-of-the-art human detector with a tracking-by-detection approach. Given these high-quality human tubes and temporal supervision, we select positive and negative tubes with very sparse spatial supervision, i.e., only one spatially annotated frame per instance. The selected tubes allow us to effectively learn a spatio-temporal action detector based on dense trajectories or CNNs. We conduct experiments on existing action localization benchmarks: UCF-Sports, J-HMDB and UCF-101. Our results show that our approach, despite using sparse spatial supervision, performs on par with methods using full supervision, i.e., one bounding box annotation per frame. To further validate our method, we introduce DALY (Daily Action Localization in YouTube), a dataset for realistic action localization in space and time. It contains high quality temporal and spatial annotations for 3.6k instances of 10 actions in 31 hours of videos (3.3M frames). It is an order of magnitude larger than existing datasets, with more diversity in appearance and long untrimmed videos.
References in corpus (6)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Weakly Supervised Object Localization with Multi-fold Multiple Instance Learning
- Learning to track for spatio-temporal action localization
- Weakly Supervised Action Labeling in Videos Under Ordering Constraints
- Action Recognition by Hierarchical Mid-level Action Elements
Cited by in corpus (19)
- ROAD: The ROad event Awareness Dataset for Autonomous Driving
- FineGym: A Hierarchical Video Dataset for Fine-grained Action Understanding
- The SARAS Endoscopic Surgeon Action Detection (ESAD) dataset: Challenges and methods
- Action Genome: Actions as Composition of Spatio-temporal Scene Graphs
- Okutama-Action: An Aerial View Video Dataset for Concurrent Human Action Detection
- Learning to Anonymize Faces for Privacy Preserving Action Detection
- Spatio-temporal Action Recognition: A Survey
- Asynchronous Temporal Fields for Action Recognition
- Two-Stream AMTnet for Action Detection
- Spatio-Temporal Action Detection with Multi-Object Interaction
- Spatio-Temporal Instance Learning: Action Tubes from Class Supervision
- Modeling Spatio-Temporal Human Track Structure for Action Localization
- Action Detection from a Robot-Car Perspective
- W-TALC: Weakly-supervised Temporal Activity Localization and Classification
- Online Spatiotemporal Action Detection and Prediction via Causal Representations
- Efficient Modelling Across Time of Human Actions and Interactions
- Discovering Multi-Label Actor-Action Association in a Weakly Supervised Setting
- TraMNet - Transition Matrix Network for Efficient Action Tube Proposals
- Hierarchical Graph-RNNs for Action Detection of Multiple Activities