activity
20172022
most citedFocal Self-attention for Local-Global Interactions in Vision Transformers

268 citations · 945 across the 30 of their papers we have counts for

collaborators

45 papers

cs.CV20221 cited

Improving Commonsense in Vision-Language Models via Knowledge Graph Riddles

Shuquan Ye, Yujia Xie, Dongdong Chen +4

This paper focuses on analyzing and improving the commonsense ability of recent popular vision-language (VL) models. Despite the great success, we observe that existing VL-models s…

cs.CV202218 cited

SinDiffusion: Learning a Diffusion Model from a Single Natural Image

Weilun Wang, Jianmin Bao, Wengang Zhou +4

We present SinDiffusion, leveraging denoising diffusion models to capture internal distribution of patches from a single natural image. SinDiffusion significantly improves the qual…

cs.CV20225 cited

Large-Scale Pre-training for Person Re-identification with Noisy Labels

Dengpan Fu, Dongdong Chen, Hao Yang +6

This paper aims to address the problem of pre-training for person re-identification (Re-ID) with noisy labels. To setup the pre-training task, we apply a simple online multi-object…

cs.CV20225 cited

Online Multi-Object Tracking with Unsupervised Re-Identification Learning and Occlusion Estimation

Qiankun Liu, Dongdong Chen, Qi Chu +4

Occlusion between different objects is a typical challenge in Multi-Object Tracking (MOT), which often leads to inferior tracking results due to the missing detected objects. The c…

cs.CV20216 cited

Unsupervised Finetuning

Suichan Li, Dongdong Chen, Yinpeng Chen +5

This paper studies "unsupervised finetuning", the symmetrical problem of the well-known "supervised finetuning". Given a pretrained model and small-scale unlabeled target data, uns…

cs.CV20213 cited

MicroNet: Improving Image Recognition with Extremely Low FLOPs

Yunsheng Li, Yinpeng Chen, Xiyang Dai +6

This paper aims at addressing the problem of substantial performance degradation at extremely low computational cost (e.g. 5M FLOPs on ImageNet classification). We found that two f…