268 citations · 377 across the 9 of their papers we have counts for
8 papers
Unified Contrastive Learning in Image-Text-Label Space
Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4
Visual recognition is recently learned via either supervised learning on human-annotated image-label data or language-image contrastive learning with webly-crawled image-text pairs…
TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment
Jianwei Yang, Yonatan Bisk, Jianfeng Gao
Contrastive learning has been widely used to train transformer-based vision-language models for video-text alignment and multi-modal representation learning. This paper presents a…
Image Scene Graph Generation (SGG) Benchmark
Xiaotian Han, Jianwei Yang, Houdong Hu +3
There is a surge of interest in image scene graph generation (object, attribute and relationship detection) due to the need of building fine-grained image understanding models that…
Focal Self-attention for Local-Global Interactions in Vision Transformers
Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4
Recently, Vision Transformer and its variants have shown great promise on various computer vision tasks. The ability of capturing short- and long-range visual dependencies through…
Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding
Pengchuan Zhang, Xiyang Dai, Jianwei Yang +4
This paper presents a new Vision Transformer (ViT) architecture Multi-Scale Vision Longformer, which significantly enhances the ViT of \cite{dosovitskiy2020image} for encoding high…
VinVL: Revisiting Visual Representations in Vision-Language Models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu +5
This paper presents a detailed study of improving visual representations for vision language (VL) tasks and develops an improved object detection model to provide object-centric re…