papers

Publications (19)

cs.CV2022

Simple Open-Vocabulary Object Detection with Vision Transformers

Matthias Minderer, Alexey Gritsenko, Austin Stone +11

cs.CV2024

PaliGemma 2: A Family of Versatile VLMs for Transfer

Andreas Steiner, André Susano Pinto, Michael Tschannen +15

cs.CV2021

On Robustness and Transferability of Convolutional Neural Networks

Josip Djolonga, Jessica Yung, Michael Tschannen +11

cs.CV2020

Unsupervised Learning of Object Structure and Dynamics from Videos

Matthias Minderer, Chen Sun, Ruben Villegas +3

cs.CV2023

Scaling Vision Transformers to 22 Billion Parameters

Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39

cs.CV2023

FlexiViT: One Model for All Patch Sizes

Lucas Beyer, Pavel Izmailov, Alexander Kolesnikov +7

cs.CV2023

Video OWL-ViT: Temporally-consistent open-world localization in video

Georg Heigold, Matthias Minderer, Alexey Gritsenko +5

cs.CV2023

Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution

Mostafa Dehghani, Basil Mustafa, Josip Djolonga +12

cs.CV2021

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov +9

cs.CV2023

PaLI-X: On Scaling up a Multilingual Vision and Language Model

Xi Chen, Josip Djolonga, Piotr Padlewski +40

cs.CV2022

Decoder Denoising Pretraining for Semantic Segmentation

Emmanuel Brempong Asiedu, Simon Kornblith, Ting Chen +3

cs.CV2021

SCENIC: A JAX Library for Computer Vision Research and Beyond

Mostafa Dehghani, Alexey Gritsenko, Anurag Arnab +2

cs.CV2020

Automatic Shortcut Removal for Self-Supervised Representation Learning

Matthias Minderer, Olivier Bachem, Neil Houlsby +1

cs.CV2024

Scaling Open-Vocabulary Object Detection

Matthias Minderer, Alexey Gritsenko, Neil Houlsby

cs.LG2025

Learning the RoPEs: Better 2D and 3D Position Encodings with STRING

Connor Schenck, Isaac Reid, Mithun George Jacob +19

cs.CV2024

Scene-Graph ViT: End-to-End Open-Vocabulary Visual Relationship Detection

Tim Salzmann, Markus Ryll, Alex Bewley +1

cs.CV2024

Improving fine-grained understanding in image-text pre-training

Ioana Bica, Anastasija Ilić, Matthias Bauer +8

cs.LG2021

Revisiting the Calibration of Modern Neural Networks

Matthias Minderer, Josip Djolonga, Rob Romijnders +5

cs.CV2024

PaliGemma: A versatile 3B VLM for transfer

Lucas Beyer, Andreas Steiner, André Susano Pinto +32