11 papers
Show, Interpret and Tell: Entity-aware Contextualised Image Captioning in Wikipedia
Khanh Nguyen, Ali Furkan Biten, Andres Mafla +2
Humans exploit prior knowledge to describe images, and are able to adapt their explanation to specific contextual information, even to the extent of inventing plausible explanation…
MUST-VQA: MUltilingual Scene-text VQA
Emanuele Vivoli, Ali Furkan Biten, Andres Mafla +2
In this paper, we present a framework for Multilingual Scene Text Visual Question Answering that deals with new languages in a zero-shot fashion. Specifically, we consider the task…
Out-of-Vocabulary Challenge Report
Sergi Garcia-Bordils, Andrés Mafla, Ali Furkan Biten +5
This paper presents final results of the Out-Of-Vocabulary 2022 (OOV) challenge. The OOV contest introduces an important aspect that is not commonly studied by Optical Character Re…
Is An Image Worth Five Sentences? A New Look into Semantics for Image-Text Matching
Ali Furkan Biten, Andres Mafla, Lluis Gomez +1
The task of image-text matching aims to map representations from different modalities into a common joint visual-textual embedding. However, the most widely used datasets for this…
StacMR: Scene-Text Aware Cross-Modal Retrieval
Andrés Mafla, Rafael Sampaio de Rezende, Lluís Gómez +2
Recent models for cross-modal retrieval have benefited from an increasingly rich understanding of visual scenes, afforded by scene graphs and object interactions to mention a few.…
Multi-Modal Reasoning Graph for Scene-Text Based Fine-Grained Image Classification and Retrieval
Andres Mafla, Sounak Dey, Ali Furkan Biten +2
Scene text instances found in natural images carry explicit semantic information that can provide important cues to solve a wide array of computer vision problems. In this paper, w…