activity
20172022
most citedFocal Self-attention for Local-Global Interactions in Vision Transformers

268 citations · 921 across the 16 of their papers we have counts for

collaborators
Showing cs.CVShow all

24 papers · 1 filter

cs.CV2022★ 67 cited

Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone

Zi-Yi Dou, Aishwarya Kamath, Zhe Gan +9

Vision-language (VL) pre-training has recently received considerable attention. However, most existing end-to-end pre-training approaches either only aim to tackle VL tasks such as…

cs.CV2022★ 126 cited

GLIPv2: Unifying Localization and Vision-Language Understanding

Haotian Zhang, Pengchuan Zhang, Xiaowei Hu +7

We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tas…

cs.CV2022

Detection Hub: Unifying Object Detection Datasets via Query Adaptation on Language Embedding

Lingchen Meng, Xiyang Dai, Yinpeng Chen +7

Combining multiple datasets enables performance boost on many computer vision tasks. But similar trend has not been witnessed in object detection when combining multiple datasets d…

cs.CV2022★ 6 cited

Unified Contrastive Learning in Image-Text-Label Space

Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4

Visual recognition is recently learned via either supervised learning on human-annotated image-label data or language-image contrastive learning with webly-crawled image-text pairs…

cs.CV2022★ 44 cited

K-LITE: Learning Transferable Visual Models with External Knowledge

Sheng Shen, Chunyuan Li, Xiaowei Hu +11

The new generation of state-of-the-art computer vision systems are trained from natural language supervision, ranging from simple object category names to descriptive captions. Thi…

cs.CV2022★ 64 cited

ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models

Chunyuan Li, Haotian Liu, Liunian Harold Li +8

Learning visual representations from natural language supervision has recently shown great promise in a number of pioneering works. In general, these language-augmented visual mode…