most citedDaViT: Dual Attention Vision Transformers

11 citations · 42 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20229 cited

Reduce Information Loss in Transformers for Pluralistic Image Inpainting

Qiankun Liu, Zhentao Tan, Dongdong Chen +6

Transformers have achieved great success in pluralistic image inpainting recently. However, we find existing transformer based solutions regard each pixel as a token, thus suffer f…

cs.CV20228 cited

Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks

Zhecan Wang, Noel Codella, Yen-Chun Chen +8

Cross-modal encoders for vision-language (VL) tasks are often pretrained with carefully curated vision-language datasets. While these datasets reach an order of 10 million samples,…

cs.CV20228 cited

MiniViT: Compressing Vision Transformers with Weight Multiplexing

Jinnian Zhang, Houwen Peng, Kan Wu +4

Vision Transformer (ViT) models have recently drawn much attention in computer vision due to their high model capability. However, ViT models suffer from huge number of parameters,…

cs.CV202211 cited

DaViT: Dual Attention Vision Transformers

Mingyu Ding, Bin Xiao, Noel Codella +3

In this work, we introduce Dual Attention Vision Transformers (DaViT), a simple yet effective vision transformer architecture that is able to capture global context while maintaini…

cs.CV20226 cited

Unified Contrastive Learning in Image-Text-Label Space

Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4

Visual recognition is recently learned via either supervised learning on human-annotated image-label data or language-image contrastive learning with webly-crawled image-text pairs…