activity
20242026
collaborators

7 papers

cs.CV2026

Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations

Jonas Klotz, Cassio F. Dantas, Pallavi Jain +2

Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy met…

cs.CV2026

RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation

Nicolas Houdré, Diego Marcos, Hugo Riffaud de Turckheim +4

Earth observation (EO) data spans a wide range of spatial, spectral, and temporal resolutions, from high-resolution optical imagery to low resolution multispectral products or rada…

cs.CV2026

Metonymy in vision models undermines attention-based interpretability

Ananthu Aniraj, Cassio F. Dantas, Dino Ienco +2

Part-based reasoning is a classical strategy to make a computer vision model directly focus on the object parts that are relevant to the downstream task. In the context of deep lea…

cs.CV2026

Two-stage Vision Transformers and Hard Masking offer Robust Object Representations

Ananthu Aniraj, Cassio F. Dantas, Dino Ienco +1

Context can strongly affect object representations, sometimes leading to undesired biases, particularly when objects appear in out-of-distribution backgrounds at inference. At the…

cs.CV2026

TimeSenCLIP: A Time Series Vision-Language Model for Remote Sensing

Pallavi Jain, Diego Marcos, Dino Ienco +2

Vision-language models (VLMs) have shown significant promise in remote sensing applications, particularly for land-use and land-cover (LULC) mapping via zero-shot classification an…

cs.CV2025

Atomizer: Generalizing to new modalities by breaking satellite images down to a set of scalars

Hugo Riffaud de Turckheim, Sylvain Lobry, Roberto Interdonato +1

The growing number of Earth observation satellites has led to increasingly diverse remote sensing data, with varying spatial, spectral, and temporal configurations. Most existing m…