22 citations · 59 across the 6 of their papers we have counts for
11 papers · 1 filter
Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos
Teng Wang, Jinrui Zhang, Feng Zheng +3
Joint video-language learning has received increasing attention in recent years. However, existing works mainly focus on single or multiple trimmed video clips (events), which make…
Poisoning MorphNet for Clean-Label Backdoor Attack to Point Clouds
Guiyu Tian, Wenhao Jiang, Wei Liu +1
This paper presents Poisoning MorphNet, the first backdoor attack method on point clouds. Conventional adversarial attack takes place in the inference stage, often fooling a model…
VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples
Tian Pan, Yibing Song, Tianyu Yang +2
MoCo is effective for unsupervised image representation learning. In this paper, we propose VideoMoCo for unsupervised video representation learning. Given a video sequence as an i…
Learning Modality Interaction for Temporal Sentence Localization and Event Captioning in Videos
Shaoxiang Chen, Wenhao Jiang, Wei Liu +1
Automatically generating sentences to describe events and temporally localizing sentences in a video are two important tasks that bridge language and videos. Recent techniques leve…
Temporally Grounding Language Queries in Videos by Contextual Boundary-aware Prediction
Jingwen Wang, Lin Ma, Wenhao Jiang
The task of temporally grounding language queries in videos is to temporally localize the best matched video segment corresponding to a given language (sentence). It requires certa…
Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network
Bairui Wang, Lin Ma, Wei Zhang +3
In this paper, we propose to guide the video caption generation with Part-of-Speech (POS) information, based on a gated fusion of multiple representations of input videos. We const…