Rethinking the Faster R-CNN Architecture for Temporal Action Localization
arXiv:1804.07667
Abstract
We propose TAL-Net, an improved approach to temporal action localization in video that is inspired by the Faster R-CNN object detection framework. TAL-Net addresses three key shortcomings of existing approaches: (1) we improve receptive field alignment using a multi-scale architecture that can accommodate extreme variation in action durations; (2) we better exploit the temporal context of actions for both proposal generation and action classification by appropriately extending receptive fields; and (3) we explicitly consider multi-stream feature fusion and demonstrate that fusing motion late is important. We achieve state-of-the-art performance for both action proposal and localization on THUMOS'14 detection benchmark and competitive performance on ActivityNet challenge.
Accepted in CVPR 2018
References in corpus (17)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- The Kinetics Human Action Video Dataset
- Multi-Scale Context Aggregation by Dilated Convolutions
- Going Deeper with Convolutions
- Rich feature hierarchies for accurate object detection and semantic segmentation
- SfM-Net: Learning of Structure and Motion from Video
- Temporal Action Localization in Untrimmed Videos via Multi-stage CNNs
- Speed/accuracy trade-offs for modern convolutional object detectors
- Learning Spatiotemporal Features with 3D Convolutional Networks
- R-C3D: Region Convolutional 3D Network for Temporal Activity Detection
- Untrimmed Video Classification for Activity Detection: submission to ActivityNet Challenge
- Temporal Context Network for Activity Localization in Videos
- TURN TAP: Temporal Unit Regression Network for Temporal Action Proposals
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- Cascaded Boundary Regression for Temporal Action Detection
- Finding Action Tubes
- Temporal Convolutional Networks for Action Segmentation and Detection