18 papers
Invisible Shortcuts: Why Vision Encoders Know Your Camera
Vladan Stojnić, Ryan Ramos, Giorgos Kordopatis-Zilos +2
Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-background or texture correl…
Benchmarking Composed Image Retrieval for Applied Earth Observation
Bill Psomas, Dionysis Christopoulos, Thanasis Petropoulos +6
Remote sensing composed image retrieval (RSCIR) enables search in large satellite image archives using composed queries that combine a reference image with a textual modifier. Alth…
Indexing Multimodal Language Models for Large-scale Image Retrieval
Bahey Tharwat, Giorgos Kordopatis-Zilos, Pavel Suma +2
Multimodal Large Language Models (MLLMs) have demonstrated strong cross-modal reasoning capabilities, yet their potential for vision-only tasks remains underexplored. We investigat…
SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation
Naomi Kombol, Ivan MartinoviÄ, SiniÅ¡a Å egviÄ +1
Foundational Vision Transformers (ViTs) have limited effectiveness in tasks requiring fine-grained spatial understanding, due to their fixed pre-training resolution and inherently…
Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
Ryan Ramos, Vladan StojniÄ, Giorgos Kordopatis-Zilos +3
Prior work has analyzed the robustness of visual encoders to image transformations and corruptions, particularly in cases where such alterations are not seen during training. When…
ELViS: Efficient Visual Similarity from Local Descriptors that Generalizes Across Domains
Pavel Suma, Giorgos Kordopatis-Zilos, Yannis Kalantidis +1
Large-scale instance-level training data is scarce, so models are typically trained on domain-specific datasets. Yet in real-world retrieval, they must handle diverse domains, maki…