papers

Publications (16)

cs.CV2022

Simple Open-Vocabulary Object Detection with Vision Transformers

Matthias Minderer, Alexey Gritsenko, Austin Stone +11

Combining simple architectures with large-scale pre-training has led to massive improvements in image classification. For object detection, pre-training and scaling approaches are…

cs.CV2019

In-domain representation learning for remote sensing

Maxim Neumann, Andre Susano Pinto, Xiaohua Zhai +1

Given the importance of remote sensing, surprisingly little attention has been paid to it by the representation learning community. To address it and to establish baselines and a c…

cs.CV2020

Training general representations for remote sensing using in-domain knowledge

Maxim Neumann, André Susano Pinto, Xiaohua Zhai +1

Automatically finding good and general remote sensing representations allows to perform transfer learning on a wide range of applications - improving the accuracy and reducing the…

cs.CV2021

Scaling Vision with Sparse Mixture of Experts

Carlos Riquelme, Joan Puigcerver, Basil Mustafa +5

Sparsely-gated Mixture of Experts networks (MoEs) have demonstrated excellent scalability in Natural Language Processing. In Computer Vision, however, almost all performant network…

cs.CV2024

Planted: a dataset for planted forest identification from multi-satellite time series

Luis Miguel Pazos-Outón, Cristina Nader Vasconcelos, Anton Raichuk +3

Protecting and restoring forest ecosystems is critical for biodiversity conservation and carbon sequestration. Forest monitoring on a global scale is essential for prioritizing and…

cs.CV2025

Uncertainty evaluation of segmentation models for Earth observation

Melanie Rey, Andriy Mnih, Maxim Neumann +2

This paper investigates methods for estimating uncertainty in semantic segmentation predictions derived from satellite imagery. Estimating uncertainty for segmentation presents uni…

cs.CV2025

Not Every Tree Is a Forest: Benchmarking Forest Types from Satellite Remote Sensing

Yuchang Jiang, Maxim Neumann

Developing accurate and reliable models for forest types mapping is critical to support efforts for halting deforestation and for biodiversity conservation (such as European Union…

cs.CV2026

Tree crop mapping of South America reveals links to deforestation and conservation

Yuchang Jiang, Anton Raichuk, Xiaoye Tong +6

Monitoring tree crop expansion is vital for zero-deforestation policies like the European Union's Regulation on Deforestation-free Products (EUDR). However, these efforts are hinde…

cs.CV2025

Zero-Shot Multi-Spectral Learning: Reimagining a Generalist Multimodal Gemini 2.5 Model for Remote Sensing Applications

Ganesh Mallya, Yotam Gigi, Dahun Kim +4

Multi-spectral imagery plays a crucial role in diverse Remote Sensing applications including land-use classification, environmental monitoring and urban planning. These images are…

cs.CV2021

Continental-Scale Building Detection from High Resolution Satellite Imagery

Wojciech Sirko, Sergii Kashubin, Marvin Ritter +7

Identifying the locations and footprints of buildings is vital for many practical and scientific purposes. Such information can be particularly useful in developing regions where a…

cs.LG2025

Heterogeneous graph neural networks for species distribution modeling

Lauren Harrell, Christine Kaeser-Chen, Burcu Karagol Ayan +7

Species distribution models (SDMs) are necessary for measuring and predicting occurrences and habitat suitability of species and their relationship with environmental factors. We i…

cs.CV2026

Reconstructing Multi-Decadal Forest Disturbances: A Spatio-Temporal Transformer Approach

Linus Scheibenreif, Anton Raichuk, Maxim Neumann

Accurate monitoring of forest disturbances is essential for understanding carbon dynamics and land management, yet traditional approaches typically rely on pixel-wise analysis of s…

cs.CV2020

AttentionNAS: Spatiotemporal Attention Cell Search for Video Classification

Xiaofang Wang, Xuehan Xiong, Maxim Neumann +5

Convolutional operations have two limitations: (1) do not explicitly model where to focus as the same filter is applied to all the positions, and (2) are unsuitable for modeling lo…

cs.CV2018

Progressive Neural Architecture Search

Chenxi Liu, Barret Zoph, Maxim Neumann +7

We propose a new method for learning the structure of convolutional neural networks (CNNs) that is more efficient than recent state-of-the-art methods based on reinforcement learni…

cs.CV2020

A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov +14

Representation learning promises to unlock deep learning for the long tail of vision tasks without expensive labelled datasets. Yet, the absence of a unified evaluation for general…

cs.CV2024

PaliGemma: A versatile 3B VLM for transfer

Lucas Beyer, Andreas Steiner, André Susano Pinto +32

PaliGemma is an open Vision-Language Model (VLM) that is based on the SigLIP-So400m vision encoder and the Gemma-2B language model. It is trained to be a versatile and broadly know…