activity
20172026
most citedMLP-Mixer: An all-MLP Architecture for Vision

1.4k citations · 2k across the 19 of their papers we have counts for

collaborators
Showing cs.CVShow all

23 papers · 1 filter

cs.CV2026

FoodSense: A Multisensory Food Dataset and Benchmark for Predicting Taste, Smell, Texture, and Sound from Images

Sabab Ishraq, Aarushi Aarushi, Juncai Jiang +1

Humans routinely infer taste, smell, texture, and even sound from food images a phenomenon well studied in cognitive science. However, prior vision language research on food has fo…

cs.CV202517 cited

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Michael Tschannen, Alexey Gritsenko, Xiao Wang +11

We introduce SigLIP 2, a family of new multilingual vision-language encoders that build on the success of the original SigLIP. In this second iteration, we extend the original imag…

cs.CV202413 cited

PaliGemma 2: A Family of Versatile VLMs for Transfer

Andreas Steiner, André Susano Pinto, Michael Tschannen +15

PaliGemma 2 is an upgrade of the PaliGemma open Vision-Language Model (VLM) based on the Gemma 2 family of language models. We combine the SigLIP-So400m vision encoder that was als…

cs.CV2024

No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language Models

Angéline Pouget, Lucas Beyer, Emanuele Bugliarello +4

We study cultural and socioeconomic diversity in contrastive vision-language models (VLMs). Using a broad range of benchmark datasets and evaluation metrics, we bring to attention…

cs.CV2024

LocCa: Visual Pretraining with Location-aware Captioners

Bo Wan, Michael Tschannen, Yongqin Xian +7

Image captioning has been shown as an effective pretraining method similar to contrastive pretraining. However, the incorporation of location-aware information into visual pretrain…

cs.CV202326 cited

PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Xi Chen, Xiao Wang, Lucas Beyer +16

This paper presents PaLI-3, a smaller, faster, and stronger vision language model (VLM) that compares favorably to similar models that are 10x larger. As part of arriving at this s…