Temporal Action Detection with Structured Segment Networks
arXiv:1704.06228
Abstract
Detecting actions in untrimmed videos is an important yet challenging task. In this paper, we present the structured segment network (SSN), a novel framework which models the temporal structure of each action instance via a structured temporal pyramid. On top of the pyramid, we further introduce a decomposed discriminative model comprising two classifiers, respectively for classifying actions and determining completeness. This allows the framework to effectively distinguish positive proposals from background or incomplete ones, thus leading to both accurate recognition and localization. These components are integrated into a unified network that can be efficiently trained in an end-to-end fashion. Additionally, a simple yet effective temporal action proposal scheme, dubbed temporal actionness grouping (TAG) is devised to generate high quality action proposals. On two challenging benchmarks, THUMOS14 and ActivityNet, our method remarkably outperforms previous state-of-the-art methods, demonstrating superior accuracy and strong adaptivity in handling actions with various temporal structures.
To appear in ICCV2017. Code & models available at http://yjxiong.me/others/ssn
References in corpus (5)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Untrimmed Video Classification for Activity Detection: submission to ActivityNet Challenge
- Temporal Activity Detection in Untrimmed Videos with Recurrent Neural Networks
- Finding Action Tubes
Cited by in corpus (30)
- Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition
- BSN: Boundary Sensitive Network for Temporal Action Proposal Generation
- Temporal Convolution Based Action Proposal: Submission to ActivityNet 2017
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language
- S3D: Single Shot multi-Span Detector via Fully 3D Convolutional Networks
- RGB-D-based Human Motion Recognition with Deep Learning: A Survey
- Step-by-step Erasion, One-by-one Collection: A Weakly Supervised Temporal Action Detector
- Local-Global Video-Text Interactions for Temporal Grounding
- Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos
- Temporal Action Proposal Generation with Transformers
- Fast Video Shot Transition Localization with Deep Structured Models
- Segregated Temporal Assembly Recurrent Networks for Weakly Supervised Multiple Action Detection
- A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer
- Move Forward and Tell: A Progressive Generator of Video Descriptions
- Temporal Human Action Segmentation via Dynamic Clustering
- Diagnosing Error in Temporal Action Detectors
- WOAD: Weakly Supervised Online Action Detection in Untrimmed Videos
- Decoupling Localization and Classification in Single Shot Temporal Action Detection
- Multilevel Language and Vision Integration for Text-to-Clip Retrieval
- On Pursuit of Designing Multi-modal Transformer for Video Grounding
- Action Search: Spotting Actions in Videos and Its Application to Temporal Action Localization
- Adversarial Seeded Sequence Growing for Weakly-Supervised Temporal Action Localization
- AutoLoc: Weakly-supervised Temporal Action Localization
- Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal Matching
- Online Detection of Action Start in Untrimmed, Streaming Videos
- Cross-modal Consensus Network for Weakly Supervised Temporal Action Localization
- Temporal Action Detection by Joint Identification-Verification
- TAN: Temporal Aggregation Network for Dense Multi-label Action Recognition
- Finding Action Tubes with a Sparse-to-Dense Framework
- Gaussian Temporal Awareness Networks for Action Localization