9 citations · 21 across the 4 of their papers we have counts for
4 papers
Masked Contrastive Pre-Training for Efficient Video-Text Retrieval
Fangxun Shu, Biaolong Chen, Yue Liao +6
We present a simple yet effective end-to-end Video-language Pre-training (VidLP) framework, Masked Contrastive Video-language Pretraining (MAC), for video-text retrieval tasks. Our…
Cross-Modality Domain Adaptation for Freespace Detection: A Simple yet Effective Baseline
Yuanbin Wang, Leyan Zhu, Shaofei Huang +4
As one of the fundamental functions of autonomous driving system, freespace detection aims at classifying each pixel of the image captured by the camera as drivable or non-drivable…
GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI Detection
Yue Liao, Aixi Zhang, Miao Lu +3
The task of Human-Object Interaction~(HOI) detection could be divided into two core problems, i.e., human-object association and interaction understanding. In this paper, we reveal…
TransRefer3D: Entity-and-Relation Aware Transformer for Fine-Grained 3D Visual Grounding
Dailan He, Yusheng Zhao, Junyu Luo +4
Recently proposed fine-grained 3D visual grounding is an essential and challenging task, whose goal is to identify the 3D object referred by a natural language sentence from other…