5 citations · 6 across the 4 of their papers we have counts for
8 papers · 1 filter
Boundary-Denoising for Video Activity Localization
Mengmeng Xu, Mattia Soldan, Jialin Gao +3
Video activity localization aims at understanding the semantic content in long untrimmed videos and retrieving actions of interest. The retrieved action with its start and end loca…
Contrastive Language-Action Pre-training for Temporal Localization
Mengmeng Xu, Erhan Gundogdu, Maksim Lapin +3
Long-form video understanding requires designing approaches that are able to temporally localize activities or language. End-to-end training for such tasks is limited by the comput…
SegTAD: Precise Temporal Action Detection via Semantic Segmentation
Chen Zhao, Merey Ramazanova, Mengmeng Xu +1
Temporal action detection (TAD) is an important yet challenging task in video analysis. Most existing works draw inspiration from image object detection and tend to reformulate it…
Low-Fidelity End-to-End Video Encoder Pre-training for Temporal Action Localization
Mengmeng Xu, Juan-Manuel Perez-Rua, Xiatian Zhu +2
Temporal action localization (TAL) is a fundamental yet challenging task in video understanding. Existing TAL methods rely on pre-training a video encoder through action classifica…
Boundary-sensitive Pre-training for Temporal Localization in Videos
Mengmeng Xu, Juan-Manuel Perez-Rua, Victor Escorcia +5
Many video analysis tasks require temporal localization thus detection of content changes. However, most existing models developed for these tasks are pre-trained on general video…
VLG-Net: Video-Language Graph Matching Network for Video Grounding
Mattia Soldan, Mengmeng Xu, Sisi Qu +2
Grounding language queries in videos aims at identifying the time interval (or moment) semantically relevant to a language query. The solution to this challenging task demands unde…