activity
20202023
most citedFocal Self-attention for Local-Global Interactions in Vision Transformers

268 citations · 377 across the 9 of their papers we have counts for

collaborators

8 papers

cs.CV20226 cited

Unified Contrastive Learning in Image-Text-Label Space

Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4

Visual recognition is recently learned via either supervised learning on human-annotated image-label data or language-image contrastive learning with webly-crawled image-text pairs…

cs.CV20214 cited

TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment

Jianwei Yang, Yonatan Bisk, Jianfeng Gao

Contrastive learning has been widely used to train transformer-based vision-language models for video-text alignment and multi-modal representation learning. This paper presents a…

cs.CV202117 cited

Image Scene Graph Generation (SGG) Benchmark

Xiaotian Han, Jianwei Yang, Houdong Hu +3

There is a surge of interest in image scene graph generation (object, attribute and relationship detection) due to the need of building fine-grained image understanding models that…

cs.CV2021268 cited

Focal Self-attention for Local-Global Interactions in Vision Transformers

Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4

Recently, Vision Transformer and its variants have shown great promise on various computer vision tasks. The ability of capturing short- and long-range visual dependencies through…

cs.CV2021

Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding

Pengchuan Zhang, Xiyang Dai, Jianwei Yang +4

This paper presents a new Vision Transformer (ViT) architecture Multi-Scale Vision Longformer, which significantly enhances the ViT of \cite{dosovitskiy2020image} for encoding high…

cs.CV202160 cited

VinVL: Revisiting Visual Representations in Vision-Language Models

Pengchuan Zhang, Xiujun Li, Xiaowei Hu +5

This paper presents a detailed study of improving visual representations for vision language (VL) tasks and develops an improved object detection model to provide object-centric re…