activity
20212024
most citedLanguage Is Not All You Need: Aligning Perception with Language Models

164 citations · 491 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2024

LADDER: An Efficient Framework for Video Frame Interpolation

Tong Shen, Dong Li, Ziheng Gao +2

Video Frame Interpolation (VFI) is a crucial technique in various applications such as slow-motion generation, frame rate conversion, video frame restoration etc. This paper introd…

cs.CV2023

Fine-grained Audible Video Description

Xuyang Shen, Dong Li, Jinxing Zhou +9

We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audibl…

cs.CV2023

Generic-to-Specific Distillation of Masked Autoencoders

Wei Huang, Zhiliang Peng, Li Dong +3

Large vision Transformers (ViTs) driven by self-supervised pre-training mechanisms achieved unprecedented progress. Lightweight ViT models limited by the model capacity, however, b…

cs.CV2022151 cited

Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Wenhui Wang, Hangbo Bao, Li Dong +8

A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves state…

cs.CV2022115 cited

BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Zhiliang Peng, Li Dong, Hangbo Bao +2

Masked image modeling (MIM) has demonstrated impressive results in self-supervised representation learning by recovering corrupted image patches. However, most existing studies ope…

cs.CV2022

HPS-Det: Dynamic Sample Assignment with Hyper-Parameter Search for Object Detection

Ji Liu, Dong Li, Zekun Li +4

Sample assignment plays a prominent part in modern object detection approaches. However, most existing methods rely on manual design to assign positive / negative samples, which do…