5 citations · 6 across the 4 of their papers we have counts for
8 papers
Contrastive Language-Action Pre-training for Temporal Localization
Mengmeng Xu, Erhan Gundogdu, Maksim Lapin +3
Long-form video understanding requires designing approaches that are able to temporally localize activities or language. End-to-end training for such tasks is limited by the comput…
SegTAD: Precise Temporal Action Detection via Semantic Segmentation
Chen Zhao, Merey Ramazanova, Mengmeng Xu +1
Temporal action detection (TAD) is an important yet challenging task in video analysis. Most existing works draw inspiration from image object detection and tend to reformulate it…
Low-Fidelity End-to-End Video Encoder Pre-training for Temporal Action Localization
Mengmeng Xu, Juan-Manuel Perez-Rua, Xiatian Zhu +2
Temporal action localization (TAL) is a fundamental yet challenging task in video understanding. Existing TAL methods rely on pre-training a video encoder through action classifica…
Boundary-sensitive Pre-training for Temporal Localization in Videos
Mengmeng Xu, Juan-Manuel Perez-Rua, Victor Escorcia +5
Many video analysis tasks require temporal localization thus detection of content changes. However, most existing models developed for these tasks are pre-trained on general video…
VLG-Net: Video-Language Graph Matching Network for Video Grounding
Mattia Soldan, Mengmeng Xu, Sisi Qu +2
Grounding language queries in videos aims at identifying the time interval (or moment) semantically relevant to a language query. The solution to this challenging task demands unde…
LC-NAS: Latency Constrained Neural Architecture Search for Point Cloud Networks
Guohao Li, Mengmeng Xu, Silvio Giancola +2
Point cloud architecture design has become a crucial problem for 3D deep learning. Several efforts exist to manually design architectures with high accuracy in point cloud tasks su…