activity
20172022
most citedDPT: Deformable Patch-based Transformer for Visual Recognition

116 citations · 301 across the 15 of their papers we have counts for

collaborators

18 papers

cs.CV20225 cited

Masked Contrastive Pre-Training for Efficient Video-Text Retrieval

Fangxun Shu, Biaolong Chen, Yue Liao +6

We present a simple yet effective end-to-end Video-language Pre-training (VidLP) framework, Masked Contrastive Video-language Pretraining (MAC), for video-text retrieval tasks. Our…

cs.CV20228 cited

Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual Tasks

Zhiyang Chen, Yousong Zhu, Zhaowen Li +8

Visual tasks vary a lot in their output formats and concerned contents, therefore it is hard to process them with an identical structure. One main obstacle lies in the high-dimensi…

cs.CV2022

UniVIP: A Unified Framework for Self-Supervised Visual Pre-training

Zhaowen Li, Yousong Zhu, Fan Yang +9

Self-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images…

cs.CV2021116 cited

DPT: Deformable Patch-based Transformer for Visual Recognition

Zhiyang Chen, Yousong Zhu, Chaoyang Zhao +4

Transformer has achieved great success in computer vision, while how to split patches in an image remains a problem. Existing methods usually use a fixed-size patch embedding which…

cs.CV202121 cited

OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation

Jing Liu, Xinxin Zhu, Fei Liu +8

In this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources. OPT is constructe…

cs.CV202129 cited

MST: Masked Self-Supervised Transformer for Visual Representation

Zhaowen Li, Zhiyang Chen, Fan Yang +8

Transformer has been widely used for self-supervised pre-training in Natural Language Processing (NLP) and achieved great success. However, it has not been fully explored in visual…