Temporal Sentence Grounding in Videos: A Survey and Future Directions
arXiv:2201.08071 · doi:10.1109/TPAMI.2023.3258628
Abstract
Temporal sentence grounding in videos (TSGV), \aka natural language video localization (NLVL) or video moment retrieval (VMR), aims to retrieve a temporal moment that semantically corresponds to a language query from an untrimmed video. Connecting computer vision and natural language, TSGV has drawn significant attention from researchers in both communities. This survey attempts to provide a summary of fundamental concepts in TSGV and current research status, as well as future research directions. As the background, we present a common structure of functional components in TSGV, in a tutorial style: from feature extraction from raw video and language query, to answer prediction of the target moment. Then we review the techniques for multimodal understanding and interaction, which is the key focus of TSGV for effective alignment between the two modalities. We construct a taxonomy of TSGV techniques and elaborate the methods in different categories with their strengths and weaknesses. Lastly, we discuss issues with the current TSGV research and share our insights about promising research directions.
Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
References in corpus (37)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Natural Language Video Localization: A Revisit in Span-based Question Answering Framework
- Video Corpus Moment Retrieval with Contrastive Learning
- VLANet: Video-Language Alignment Network for Weakly-Supervised Video Moment Retrieval
- ExCL: Extractive Clip Localization Using Natural Language Descriptions
- Frame-wise Cross-modal Matching for Video Moment Retrieval
- A Closer Look at Temporal Sentence Grounding in Videos: Dataset and Metric
- Look Closer to Ground Better: Weakly-Supervised Temporal Grounding of Sentence in Video
- Multi-Modal Interaction Graph Convolutional Network for Temporal Language Localization in Videos
- Parallel Attention Network with Sequence Matching for Video Grounding
- CONQUER: Contextual Query-aware Ranking for Video Corpus Moment Retrieval
- Video Moment Retrieval from Text Queries via Single Frame Annotation
- You Need to Read Again: Multi-granularity Perception Network for Moment Retrieval in Videos
- Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
- Text-based Localization of Moments in a Video Corpus
- QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries
- End-to-end Multi-modal Video Temporal Grounding
- A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus
- Fine-grained Semantic Alignment Network for Weakly Supervised Temporal Language Grounding
- Deep Graph Random Process for Relational-Thinking-Based Speech Recognition
- Hierarchical Deep Residual Reasoning for Temporal Moment Localization
- Towards Debiasing Temporal Sentence Grounding in Video
- End-to-End Dense Video Grounding via Parallel Regression
- A Simple Yet Effective Method for Video Temporal Grounding with Cross-Modality Attention
- A Survey on Natural Language Video Localization
- Contrastive Language-Action Pre-training for Temporal Localization
- Multi-Scale Self-Contrastive Learning with Hard Negative Mining for Weakly-Supervised Query-based Video Grounding
- Explore-And-Match: Bridging Proposal-Based and Proposal-Free With Transformer for Sentence Grounding in Videos
- LocFormer: Enabling Transformers to Perform Temporal Moment Localization on Long Untrimmed Videos With a Feature Sampling Approach
- Decoupled Spatial Temporal Graphs for Generic Visual Grounding
- SNEAK: Synonymous Sentences-Aware Adversarial Attack on Natural Language Video Localization
- A Multi-level Alignment Training Scheme for Video-and-Language Grounding
- Learning Commonsense-aware Moment-Text Alignment for Fast Video Temporal Grounding
- Video Moment Retrieval with Text Query Considering Many-to-Many Correspondence Using Potentially Relevant Pair
- Entity-aware and Motion-aware Transformers for Language-driven Action Localization in Videos
- Team PKU-WICT-MIPL PIC Makeup Temporal Video Grounding Challenge 2022 Technical Report