6 papers
COP-GEN: Latent Diffusion Transformer for Copernicus Earth Observation Data
Miguel Espinosa, Eva Gmelich Meijling, Valerio Marsocci +2
Earth observation applications increasingly rely on data from multiple sensors, including optical, radar, elevation, and land-cover. Relationships between modalities are fundamenta…
No time to train! Training-Free Reference-Based Instance Segmentation
Miguel Espinosa, Chenhongyi Yang, Linus Ericsson +2
The performance of image segmentation models has historically been constrained by the high cost of collecting large-scale annotated data. The Segment Anything Model (SAM) alleviate…
COP-GEN-Beta: Unified Generative Modelling of COPernicus Imagery Thumbnails
Miguel Espinosa, Valerio Marsocci, Yuru Jia +2
In remote sensing, multi-modal data from various sensors capturing the same scene offers rich opportunities, but learning a unified representation across these modalities remains a…
There is no SAMantics! Exploring SAM as a Backbone for Visual Understanding Tasks
Miguel Espinosa, Chenhongyi Yang, Linus Ericsson +2
The Segment Anything Model (SAM) was originally designed for label-agnostic mask generation. Does this model also possess inherent semantic understanding, of value to broader visua…
einspace: Searching for Neural Architectures from Fundamental Operations
Linus Ericsson, Miguel Espinosa, Chenhongyi Yang +5
Neural architecture search (NAS) finds high performing networks for a given task. Yet the results of NAS are fairly prosaic; they did not e.g. create a shift from convolutional str…
PlainMamba: Improving Non-Hierarchical Mamba in Visual Recognition
Chenhongyi Yang, Zehui Chen, Miguel Espinosa +4
We present PlainMamba: a simple non-hierarchical state space model (SSM) designed for general visual recognition. The recent Mamba model has shown how SSMs can be highly competitiv…