activity
20182021
collaborators

11 papers

cs.CV20222 cited

Show, Interpret and Tell: Entity-aware Contextualised Image Captioning in Wikipedia

Khanh Nguyen, Ali Furkan Biten, Andres Mafla +2

Humans exploit prior knowledge to describe images, and are able to adapt their explanation to specific contextual information, even to the extent of inventing plausible explanation…

cs.CV20221 cited

MUST-VQA: MUltilingual Scene-text VQA

Emanuele Vivoli, Ali Furkan Biten, Andres Mafla +2

In this paper, we present a framework for Multilingual Scene Text Visual Question Answering that deals with new languages in a zero-shot fashion. Specifically, we consider the task…

cs.CV2022

Out-of-Vocabulary Challenge Report

Sergi Garcia-Bordils, Andrés Mafla, Ali Furkan Biten +5

This paper presents final results of the Out-Of-Vocabulary 2022 (OOV) challenge. The OOV contest introduces an important aspect that is not commonly studied by Optical Character Re…

cs.CV2021

Is An Image Worth Five Sentences? A New Look into Semantics for Image-Text Matching

Ali Furkan Biten, Andres Mafla, Lluis Gomez +1

The task of image-text matching aims to map representations from different modalities into a common joint visual-textual embedding. However, the most widely used datasets for this…

cs.CV2020

StacMR: Scene-Text Aware Cross-Modal Retrieval

Andrés Mafla, Rafael Sampaio de Rezende, Lluís Gómez +2

Recent models for cross-modal retrieval have benefited from an increasingly rich understanding of visual scenes, afforded by scene graphs and object interactions to mention a few.…

cs.CV2020

Multi-Modal Reasoning Graph for Scene-Text Based Fine-Grained Image Classification and Retrieval

Andres Mafla, Sounak Dey, Ali Furkan Biten +2

Scene text instances found in natural images carry explicit semantic information that can provide important cues to solve a wide array of computer vision problems. In this paper, w…