268 citations · 631 across the 13 of their papers we have counts for
11 papers · 1 filter
Semantic Cross Attention for Few-shot Learning
Bin Xiao, Chien-Liang Liu, Wen-Hoar Hsaio
Few-shot learning (FSL) has attracted considerable attention recently. Among existing approaches, the metric-based method aims to train an embedding network that can make similar s…
Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks
Zhecan Wang, Noel Codella, Yen-Chun Chen +8
Cross-modal encoders for vision-language (VL) tasks are often pretrained with carefully curated vision-language datasets. While these datasets reach an order of 10 million samples,…
MiniViT: Compressing Vision Transformers with Weight Multiplexing
Jinnian Zhang, Houwen Peng, Kan Wu +4
Vision Transformer (ViT) models have recently drawn much attention in computer vision due to their high model capability. However, ViT models suffer from huge number of parameters,…
DaViT: Dual Attention Vision Transformers
Mingyu Ding, Bin Xiao, Noel Codella +3
In this work, we introduce Dual Attention Vision Transformers (DaViT), a simple yet effective vision transformer architecture that is able to capture global context while maintaini…
Unified Contrastive Learning in Image-Text-Label Space
Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4
Visual recognition is recently learned via either supervised learning on human-annotated image-label data or language-image contrastive learning with webly-crawled image-text pairs…
Focal Self-attention for Local-Global Interactions in Vision Transformers
Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4
Recently, Vision Transformer and its variants have shown great promise on various computer vision tasks. The ability of capturing short- and long-range visual dependencies through…