TURN TAP: Temporal Unit Regression Network for Temporal Action Proposals
arXiv:1703.06189
Abstract
Temporal Action Proposal (TAP) generation is an important problem, as fast and accurate extraction of semantically important (e.g. human actions) segments from untrimmed videos is an important step for large-scale video analysis. We propose a novel Temporal Unit Regression Network (TURN) model. There are two salient aspects of TURN: (1) TURN jointly predicts action proposals and refines the temporal boundaries by temporal coordinate regression; (2) Fast computation is enabled by unit feature reuse: a long untrimmed video is decomposed into video units, which are reused as basic building blocks of temporal proposals. TURN outperforms the state-of-the-art methods under average recall (AR) by a large margin on THUMOS-14 and ActivityNet datasets, and runs at over 880 frames per second (FPS) on a TITAN X GPU. We further apply TURN as a proposal generation stage for existing temporal action localization pipelines, it outperforms state-of-the-art performance on THUMOS-14 and ActivityNet.
ICCV 2017 camera ready
Cited by in corpus (26)
- End-to-End Dense Video Captioning with Masked Transformer
- Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning
- Jointly Localizing and Describing Events for Dense Video Captioning
- Rethinking the Faster R-CNN Architecture for Temporal Action Localization
- RGB-D-based Human Motion Recognition with Deep Learning: A Survey
- Step-by-step Erasion, One-by-one Collection: A Weakly Supervised Temporal Action Detector
- To Find Where You Talk: Temporal Sentence Localization in Video with Attention Based Location Regression
- Motion-Appearance Co-Memory Networks for Video Question Answering
- Multi-granularity Generator for Temporal Action Proposal
- Temporal Action Proposal Generation with Transformers
- Long Short-Term Transformer for Online Action Detection
- Fully-Coupled Two-Stream Spatiotemporal Networks for Extremely Low Resolution Action Recognition
- Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences
- Diagnosing Error in Temporal Action Detectors
- Low-Fidelity End-to-End Video Encoder Pre-training for Temporal Action Localization
- Decoupling Localization and Classification in Single Shot Temporal Action Detection
- Action Search: Spotting Actions in Videos and Its Application to Temporal Action Localization
- AutoLoc: Weakly-supervised Temporal Action Localization
- Follow the Attention: Combining Partial Pose and Object Motion for Fine-Grained Action Detection
- Localizing the Common Action Among a Few Videos
- Online Detection of Action Start in Untrimmed, Streaming Videos
- Temporal Action Localization using Long Short-Term Dependency
- DORi: Discovering Object Relationship for Moment Localization of a Natural-Language Query in Video
- Gaussian Temporal Awareness Networks for Action Localization
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions
- Data-efficient Alignment of Multimodal Sequences by Aligning Gradient Updates and Internal Feature Distributions