activity
20242026
collaborators

5 papers

cs.CV2026

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference

Imanol Miranda, Ander Salaberria, Eneko Agirre +1

Dual-encoder Vision-Language Models (VLMs) such as CLIP are often characterized as bag-of-words systems due to their poor performance on compositional benchmarks. We argue that thi…

cs.CL2026

Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque

Lukas Arana, Julen Etxaniz, Ander Salaberria +1

Current Multimodal Large Language Models exhibit very strong performance for several demanding tasks. While commercial MLLMs deliver acceptable performance in low-resource language…

cs.CV2025

Adding simple structure at inference improves Vision-Language Compositionality

Imanol Miranda, Ander Salaberria, Eneko Agirre +1

Dual encoder Vision-Language Models (VLM) such as CLIP are widely used for image-text retrieval tasks. However, those models struggle with compositionality, showing a bag-of-words-…

cs.CL2025

Vision-Language Models Struggle to Align Entities across Modalities

Iñigo Alonso, Gorka Azkune, Ander Salaberria +2

Cross-modal entity linking refers to the ability to align entities and their attributes across different modalities. While cross-modal entity linking is a fundamental skill needed…

cs.CV2024

BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval

Imanol Miranda, Ander Salaberria, Eneko Agirre +1

Existing Vision-Language Compositionality (VLC) benchmarks like SugarCrepe are formulated as image-to-text retrieval problems, where, given an image, the models need to select betw…