A Survey on Natural Language Video Localization
arXiv:2104.00234
Abstract
Natural language video localization (NLVL), which aims to locate a target moment from a video that semantically corresponds to a text query, is a novel and challenging task. Toward this end, in this paper, we present a comprehensive survey of the NLVL algorithms, where we first propose the pipeline of NLVL, and then categorize them into supervised and weakly-supervised methods, following by the analysis of the strengths and weaknesses of each kind of methods. Subsequently, we present the dataset, evaluation protocols and the general performance analysis. Finally, the possible perspectives are obtained by summarizing the existing methods.
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- ExCL: Extractive Clip Localization Using Natural Language Descriptions
- Local-Global Video-Text Interactions for Temporal Grounding
- Dense Regression Network for Video Grounding
- Learning Modality Interaction for Temporal Sentence Localization and Event Captioning in Videos
- Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed Videos
- Tree-Structured Policy based Progressive Reinforcement Learning for Temporally Language Grounding in Video