activity
20172022
most citedFocal Self-attention for Local-Global Interactions in Vision Transformers

268 citations · 596 across the 9 of their papers we have counts for

collaborators

15 papers

cs.CV20226 cited

Unified Contrastive Learning in Image-Text-Label Space

Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4

Visual recognition is recently learned via either supervised learning on human-annotated image-label data or language-image contrastive learning with webly-crawled image-text pairs…

cs.CV202230 cited

Vision-Language Intelligence: Tasks, Representation Learning, and Large Models

Feng Li, Hao Zhang, Yi-Fan Zhang +5

This paper presents a comprehensive survey of vision-language (VL) intelligence from the perspective of time. This survey is inspired by the remarkable progress in both computer vi…

cs.CV202117 cited

Image Scene Graph Generation (SGG) Benchmark

Xiaotian Han, Jianwei Yang, Houdong Hu +3

There is a surge of interest in image scene graph generation (object, attribute and relationship detection) due to the need of building fine-grained image understanding models that…

cs.CV2021268 cited

Focal Self-attention for Local-Global Interactions in Vision Transformers

Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4

Recently, Vision Transformer and its variants have shown great promise on various computer vision tasks. The ability of capturing short- and long-range visual dependencies through…

cs.CV2021

Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding

Pengchuan Zhang, Xiyang Dai, Jianwei Yang +4

This paper presents a new Vision Transformer (ViT) architecture Multi-Scale Vision Longformer, which significantly enhances the ViT of \cite{dosovitskiy2020image} for encoding high…

cs.CV202160 cited

VinVL: Revisiting Visual Representations in Vision-Language Models

Pengchuan Zhang, Xiujun Li, Xiaowei Hu +5

This paper presents a detailed study of improving visual representations for vision language (VL) tasks and develops an improved object detection model to provide object-centric re…