activity
20212023
most citedImage as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

151 citations · 327 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CV2023

Fine-grained Audible Video Description

Xuyang Shen, Dong Li, Jinxing Zhou +9

We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audibl…

cs.CL2023164 cited

Language Is Not All You Need: Aligning Perception with Language Models

Shaohan Huang, Li Dong, Wenhui Wang +15

A big convergence of language, multimodal perception, action, and world modeling is a key step toward artificial general intelligence. In this work, we introduce Kosmos-1, a Multim…

cs.CV2023

Generic-to-Specific Distillation of Masked Autoencoders

Wei Huang, Zhiliang Peng, Li Dong +3

Large vision Transformers (ViTs) driven by self-supervised pre-training mechanisms achieved unprecedented progress. Lightweight ViT models limited by the model capacity, however, b…

cs.CV2022151 cited

Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Wenhui Wang, Hangbo Bao, Li Dong +8

A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves state…

cs.CV2022115 cited

BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Zhiliang Peng, Li Dong, Hangbo Bao +2

Masked image modeling (MIM) has demonstrated impressive results in self-supervised representation learning by recovering corrupted image patches. However, most existing studies ope…

cs.CV2022

HPS-Det: Dynamic Sample Assignment with Hyper-Parameter Search for Object Detection

Ji Liu, Dong Li, Zekun Li +4

Sample assignment plays a prominent part in modern object detection approaches. However, most existing methods rely on manual design to assign positive / negative samples, which do…