1k citations · 1.3k across the 8 of their papers we have counts for
Showing 2022 · cs.CVShow all
2 papers · 2 filters
cs.CV2022★ 8 cited
Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks
Zhecan Wang, Noel Codella, Yen-Chun Chen +8
Cross-modal encoders for vision-language (VL) tasks are often pretrained with carefully curated vision-language datasets. While these datasets reach an order of 10 million samples,…
cs.CV2022★ 11 cited
DaViT: Dual Attention Vision Transformers
Mingyu Ding, Bin Xiao, Noel Codella +3
In this work, we introduce Dual Attention Vision Transformers (DaViT), a simple yet effective vision transformer architecture that is able to capture global context while maintaini…