activity
20172023
most citedFocal Self-attention for Local-Global Interactions in Vision Transformers

268 citations · 952 across the 33 of their papers we have counts for

collaborators
Showing cs.CVShow all

48 papers · 1 filter

cs.CV2023

Improving Adversarial Robustness of Masked Autoencoders via Test-time Frequency-domain Prompting

Qidong Huang, Xiaoyi Dong, Dongdong Chen +5

In this paper, we investigate the adversarial robustness of vision transformers that are equipped with BERT pretraining (e.g., BEiT, MAE). A surprising observation is that MAE has…

cs.CV20232 cited

HQ-50K: A Large-scale, High-quality Dataset for Image Restoration

Qinhong Yang, Dongdong Chen, Zhentao Tan +6

This paper introduces a new large-scale image restoration dataset, called HQ-50K, which contains 50,000 high-quality images with rich texture details and semantic diversity. We ana…

cs.CV20235 cited

Designing a Better Asymmetric VQGAN for StableDiffusion

Zixin Zhu, Xuelu Feng, Dongdong Chen +5

StableDiffusion is a revolutionary text-to-image generator that is causing a stir in the world of image generation and editing. Unlike traditional methods that learn a diffusion mo…

cs.CV2023

Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models

Shihao Zhao, Dongdong Chen, Yen-Chun Chen +4

Text-to-Image diffusion models have made tremendous progress over the past two years, enabling the generation of highly realistic images based on open-domain text descriptions. How…

cs.CV20221 cited

Improving Commonsense in Vision-Language Models via Knowledge Graph Riddles

Shuquan Ye, Yujia Xie, Dongdong Chen +4

This paper focuses on analyzing and improving the commonsense ability of recent popular vision-language (VL) models. Despite the great success, we observe that existing VL-models s…

cs.CV202218 cited

SinDiffusion: Learning a Diffusion Model from a Single Natural Image

Weilun Wang, Jianmin Bao, Wengang Zhou +4

We present SinDiffusion, leveraging denoising diffusion models to capture internal distribution of patches from a single natural image. SinDiffusion significantly improves the qual…