activity
20212023
most citedCoCa: Contrastive Captioners are Image-Text Foundation Models

518 citations · 883 across the 6 of their papers we have counts for

collaborators

8 papers

cs.CV2023★ 4 cited

CoBIT: A Contrastive Bi-directional Image-Text Generation Model

Haoxuan You, Mandy Guo, Zhecan Wang +3

The field of vision and language has witnessed a proliferation of pre-trained foundation models. Most existing methods are independently pre-trained with contrastive objective like…

cs.CV2022★ 20 cited

VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Shen Yan, Tao Zhu, Zirui Wang +5

We explore an efficient approach to establish a foundational video-text model. We present VideoCoCa that maximally reuses a pretrained image-text contrastive captioner (CoCa) model…

cs.CV2022★ 1 cited

Exploiting Category Names for Few-Shot Classification with Vision-Language Models

Taihong Xiao, Zirui Wang, Liangliang Cao +3

Vision-language foundation models pretrained on large-scale data provide a powerful tool for many visual understanding tasks. Notably, many vision-language models build two encoder…

cs.CV2022★ 340 cited

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Jiahui Yu, Yuanzhong Xu, Jing Yu Koh +14

We present the Pathways Autoregressive Text-to-Image (Parti) model, which generates high-fidelity photorealistic images and supports content-rich synthesis involving complex compos…

cs.CV2022★ 518 cited

CoCa: Contrastive Captioners are Image-Text Foundation Models

Jiahui Yu, Zirui Wang, Vijay Vasudevan +3

Exploring large-scale pretrained foundation models is of significant interest in computer vision because these models can be quickly transferred to many downstream tasks. This pape…

cs.LG2021

Combined Scaling for Zero-shot Transfer Learning

Hieu Pham, Zihang Dai, Golnaz Ghiasi +9

We present a combined scaling method - named BASIC - that achieves 85.7% top-1 accuracy on the ImageNet ILSVRC-2012 validation set without learning from any labeled ImageNet exampl…