160 citations · 351 across the 9 of their papers we have counts for
7 papers · 1 filter
Jet: A Modern Transformer-Based Normalizing Flow
Alexander Kolesnikov, André Susano Pinto, Michael Tschannen
In the past, normalizing generative flows have emerged as a promising class of generative models for natural images. This type of model has many modeling advantages: the ability to…
PaliGemma 2: A Family of Versatile VLMs for Transfer
Andreas Steiner, André Susano Pinto, Michael Tschannen +15
PaliGemma 2 is an upgrade of the PaliGemma open Vision-Language Model (VLM) based on the Gemma 2 family of language models. We combine the SigLIP-So400m vision encoder that was als…
Learning to Merge Tokens in Vision Transformers
Cedric Renggli, André Susano Pinto, Neil Houlsby +3
Transformers are widely applied to solve natural language understanding and computer vision tasks. While scaling up these architectures leads to improved performance, it often come…
Scaling Vision with Sparse Mixture of Experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa +5
Sparsely-gated Mixture of Experts networks (MoEs) have demonstrated excellent scalability in Natural Language Processing. In Computer Vision, however, almost all performant network…
Training general representations for remote sensing using in-domain knowledge
Maxim Neumann, André Susano Pinto, Xiaohua Zhai +1
Automatically finding good and general remote sensing representations allows to perform transfer learning on a wide range of applications - improving the accuracy and reducing the…
In-domain representation learning for remote sensing
Maxim Neumann, Andre Susano Pinto, Xiaohua Zhai +1
Given the importance of remote sensing, surprisingly little attention has been paid to it by the representation learning community. To address it and to establish baselines and a c…