Temporal Context Network for Activity Localization in Videos
arXiv:1708.02349
Abstract
We present a Temporal Context Network (TCN) for precise temporal localization of human activities. Similar to the Faster-RCNN architecture, proposals are placed at equal intervals in a video which span multiple temporal scales. We propose a novel representation for ranking these proposals. Since pooling features only inside a segment is not sufficient to predict activity boundaries, we construct a representation which explicitly captures context around a proposal for ranking it. For each temporal segment inside a proposal, features are uniformly sampled at a pair of scales and are input to a temporal convolutional neural network for classification. After ranking proposals, non-maximum suppression is applied and classification is performed to obtain final detections. TCN outperforms state-of-the-art methods on the ActivityNet dataset and the THUMOS14 dataset.
To appear in ICCV 2017
References in corpus (4)
Cited by in corpus (6)
- Contextual Multi-Scale Region Convolutional 3D Network for Activity Detection
- Decoupling Localization and Classification in Single Shot Temporal Action Detection
- Exploring Uncertainty in Conditional Multi-Modal Retrieval Systems
- Follow the Attention: Combining Partial Pose and Object Motion for Fine-Grained Action Detection
- Localizing the Common Action Among a Few Videos
- TAN: Temporal Aggregation Network for Dense Multi-label Action Recognition