Temporal Action Localization in Untrimmed Videos via Multi-stage CNNs
arXiv:1601.02129
Abstract
We address temporal action localization in untrimmed long videos. This is important because videos in real applications are usually unconstrained and contain multiple action instances plus video content of background scenes or other activities. To address this challenging issue, we exploit the effectiveness of deep networks in temporal action localization via three segment-based 3D ConvNets: (1) a proposal network identifies candidate segments in a long video that may contain actions; (2) a classification network learns one-vs-all action classification model to serve as initialization for the localization network; and (3) a localization network fine-tunes on the learned classification network to localize each action instance. We propose a novel loss function for the localization network to explicitly consider temporal overlap and therefore achieve high temporal localization accuracy. Only the proposal network and the localization network are used during prediction. On two large-scale benchmarks, our approach achieves significantly superior performances compared with other state-of-the-art systems: mAP increases from 1.7% to 7.4% on MEXaction2 and increases from 15.0% to 19.0% on THUMOS 2014, when the overlap threshold for evaluation is set to 0.5.
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016
Cited by in corpus (54)
- The THUMOS Challenge on Action Recognition for Videos "in the Wild"
- A Closer Look at Spatiotemporal Convolutions for Action Recognition
- Temporal Action Detection with Structured Segment Networks
- A Pursuit of Temporal Accuracy in General Activity Detection
- R-C3D: Region Convolutional 3D Network for Temporal Activity Detection
- Temporal Activity Detection in Untrimmed Videos with Recurrent Neural Networks
- Temporal Context Network for Activity Localization in Videos
- Temporal Convolution Based Action Proposal: Submission to ActivityNet 2017
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- Cascaded Boundary Regression for Temporal Action Detection
- Weakly Supervised Action Localization by Sparse Temporal Pooling Network
- S3D: Single Shot multi-Span Detector via Fully 3D Convolutional Networks
- End-to-End Dense Video Captioning with Masked Transformer
- Predictive-Corrective Networks for Action Detection
- Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning
- Tripping through time: Efficient Localization of Activities in Videos
- Jointly Localizing and Describing Events for Dense Video Captioning
- Rethinking the Faster R-CNN Architecture for Temporal Action Localization
- UntrimmedNets for Weakly Supervised Action Recognition and Detection
- Local-Global Video-Text Interactions for Temporal Grounding
- RED: Reinforced Encoder-Decoder Networks for Action Anticipation
- Motion-Appearance Co-Memory Networks for Video Question Answering
- Multi-granularity Generator for Temporal Action Proposal
- Temporal Action Proposal Generation with Transformers
- Contextual Multi-Scale Region Convolutional 3D Network for Activity Detection
- Spatio-Temporal Action Detection with Cascade Proposal and Location Anticipation
- Long Short-Term Transformer for Online Action Detection
- DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition
- Segregated Temporal Assembly Recurrent Networks for Weakly Supervised Multiple Action Detection
- Localizing Unseen Activities in Video via Image Query
- Action Completion: A Temporal Model for Moment Detection
- A Survey on Natural Language Video Localization
- WOAD: Weakly Supervised Online Action Detection in Untrimmed Videos
- Low-Fidelity End-to-End Video Encoder Pre-training for Temporal Action Localization
- Temporal Tessellation: A Unified Approach for Video Analysis
- A Self-Adaptive Proposal Model for Temporal Action Detection based on Reinforcement Learning
- Budget-Aware Activity Detection with A Recurrent Policy Network
- Context-aware and Scale-insensitive Temporal Repetition Counting
- Action Sets: Weakly Supervised Action Segmentation without Ordering Constraints
- On Pursuit of Designing Multi-modal Transformer for Video Grounding
- Graph Distillation for Action Detection with Privileged Modalities
- Efficient Action Detection in Untrimmed Videos via Multi-Task Learning
- Improved Soccer Action Spotting using both Audio and Video Streams
- Cross-modal Consensus Network for Weakly Supervised Temporal Action Localization
- Action Understanding with Multiple Classes of Actors
- Localizing the Common Action Among a Few Videos
- A Comprehensive Study on Temporal Modeling for Online Action Detection
- Gaussian Temporal Awareness Networks for Action Localization
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions
- TAN: Temporal Aggregation Network for Dense Multi-label Action Recognition
- Language Guided Networks for Cross-modal Moment Retrieval
- Temporal Action Detection by Joint Identification-Verification
- Accurate Temporal Action Proposal Generation with Relation-Aware Pyramid Network
- Towards Diverse Paragraph Captioning for Untrimmed Videos