activity
20182023
most citedVideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples

22 citations · 59 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2023★ 7 cited

Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos

Teng Wang, Jinrui Zhang, Feng Zheng +3

Joint video-language learning has received increasing attention in recent years. However, existing works mainly focus on single or multiple trimmed video clips (events), which make…

cs.CV2021

Poisoning MorphNet for Clean-Label Backdoor Attack to Point Clouds

Guiyu Tian, Wenhao Jiang, Wei Liu +1

This paper presents Poisoning MorphNet, the first backdoor attack method on point clouds. Conventional adversarial attack takes place in the inference stage, often fooling a model…

cs.CV2021★ 22 cited

VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples

Tian Pan, Yibing Song, Tianyu Yang +2

MoCo is effective for unsupervised image representation learning. In this paper, we propose VideoMoCo for unsupervised video representation learning. Given a video sequence as an i…

cs.CV2020★ 11 cited

Learning Modality Interaction for Temporal Sentence Localization and Event Captioning in Videos

Shaoxiang Chen, Wenhao Jiang, Wei Liu +1

Automatically generating sentences to describe events and temporally localizing sentences in a video are two important tasks that bridge language and videos. Recent techniques leve…

cs.CV2019

Temporally Grounding Language Queries in Videos by Contextual Boundary-aware Prediction

Jingwen Wang, Lin Ma, Wenhao Jiang

The task of temporally grounding language queries in videos is to temporally localize the best matched video segment corresponding to a given language (sentence). It requires certa…

cs.CV2019

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network

Bairui Wang, Lin Ma, Wei Zhang +3

In this paper, we propose to guide the video caption generation with Part-of-Speech (POS) information, based on a gated fusion of multiple representations of input videos. We const…