3 papers
cs.CV2026
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
Imanol Miranda, Ander Salaberria, Eneko Agirre +1
Dual-encoder Vision-Language Models (VLMs) such as CLIP are often characterized as bag-of-words systems due to their poor performance on compositional benchmarks. We argue that thi…
cs.CV2025
TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
Iñigo Alonso, Imanol Miranda, Eneko Agirre +1
While table understanding increasingly relies on pixel-only settings, current benchmarks predominantly use synthetic renderings that lack the complexity and visual diversity of rea…
cs.CV2025
Adding simple structure at inference improves Vision-Language Compositionality
Imanol Miranda, Ander Salaberria, Eneko Agirre +1
Dual encoder Vision-Language Models (VLM) such as CLIP are widely used for image-text retrieval tasks. However, those models struggle with compositionality, showing a bag-of-words-…