activity
20152024
most citedUniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training

225 citations · 1.4k across the 72 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2023

TextDiffuser: Diffusion Models as Text Painters

Jingye Chen, Yupan Huang, Tengchao Lv +3

Diffusion models have gained increasing attention for their impressive generation abilities but currently struggle with rendering accurate and coherent text. To address this issue,…

cs.CV202214 cited

A Unified View of Masked Image Modeling

Zhiliang Peng, Li Dong, Hangbo Bao +2

Masked image modeling has demonstrated great potential to eliminate the label-hungry problem of training large-scale vision Transformers, achieving impressive performance on variou…

cs.CV20222 cited

Non-Contrastive Learning Meets Language-Image Pre-Training

Jinghao Zhou, Li Dong, Zhe Gan +2

Contrastive language-image pre-training (CLIP) serves as a de-facto standard to align images and texts. Nonetheless, the loose correlation between images and texts of web-crawled d…

cs.CV20222 cited

CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Haoyu Song, Li Dong, Wei-Nan Zhang +2

CLIP has shown a remarkable zero-shot capability on a wide range of vision tasks. Previously, CLIP is only regarded as a powerful visual encoder. However, after being pre-trained b…

cs.CV2020

Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks

Xiujun Li, Xi Yin, Chunyuan Li +9

Large-scale pre-training methods of learning cross-modal representations on image-text pairs are becoming popular for vision-language tasks. While existing methods simply concatena…

cs.CV2019

VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Weijie Su, Xizhou Zhu, Yue Cao +4

We introduce a new pre-trainable generic representation for visual-linguistic tasks, called Visual-Linguistic BERT (VL-BERT for short). VL-BERT adopts the simple yet powerful Trans…