activity
20152026
most citedMLP-Mixer: An all-MLP Architecture for Vision

1.4k citations · 2.1k across the 20 of their papers we have counts for

collaborators
Showing 2023 · cs.CVShow all

7 papers · 2 filters

cs.CV2023★ 26 cited

PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Xi Chen, Xiao Wang, Lucas Beyer +16

This paper presents PaLI-3, a smaller, faster, and stronger vision language model (VLM) that compares favorably to similar models that are 10x larger. As part of arriving at this s…

cs.CV2023★ 39 cited

PaLI-X: On Scaling up a Multilingual Vision and Language Model

Xi Chen, Josip Djolonga, Piotr Padlewski +40

We present the training recipe and results of scaling up PaLI-X, a multilingual vision and language model, both in terms of size of the components and the breadth of its training t…

cs.CV2023★ 4 cited

Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design

Ibrahim Alabdulmohsin, Xiaohua Zhai, Alexander Kolesnikov +1

Scaling laws have been recently employed to derive compute-optimal model size (number of parameters) for a given compute duration. We advance and refine such methods to infer compu…

cs.CV2023★ 3 cited

A Study of Autoregressive Decoders for Multi-Tasking in Computer Vision

Lucas Beyer, Bo Wan, Gagan Madan +9

There has been a recent explosion of computer vision models which perform many tasks and are composed of an image encoder (usually a ViT) and an autoregressive decoder (usually a T…

cs.CV2023★ 17 cited

Sigmoid Loss for Language Image Pre-Training

Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov +1

We propose a simple pairwise Sigmoid loss for Language-Image Pre-training (SigLIP). Unlike standard contrastive learning with softmax normalization, the sigmoid loss operates solel…

cs.CV2023★ 13 cited

Tuning computer vision models with task rewards

André Susano Pinto, Alexander Kolesnikov, Yuge Shi +2

Misalignment between model predictions and intended usage can be detrimental for the deployment of computer vision models. The issue is exacerbated when the task involves complex s…