12 citations · 20 across the 13 of their papers we have counts for
24 papers
Show, Interpret and Tell: Entity-aware Contextualised Image Captioning in Wikipedia
Khanh Nguyen, Ali Furkan Biten, Andres Mafla +2
Humans exploit prior knowledge to describe images, and are able to adapt their explanation to specific contextual information, even to the extent of inventing plausible explanation…
MUST-VQA: MUltilingual Scene-text VQA
Emanuele Vivoli, Ali Furkan Biten, Andres Mafla +2
In this paper, we present a framework for Multilingual Scene Text Visual Question Answering that deals with new languages in a zero-shot fashion. Specifically, we consider the task…
A Generic Image Retrieval Method for Date Estimation of Historical Document Collections
Adrià Molina, Lluis Gomez, Oriol Ramos Terrades +1
Date estimation of historical document images is a challenging problem, with several contributions in the literature that lack of the ability to generalize from one dataset to othe…
OCR-IDL: OCR Annotations for Industry Document Library Dataset
Ali Furkan Biten, Rubèn Tito, Lluis Gomez +2
Pretraining has proven successful in Document Intelligence tasks where deluge of documents are used to pretrain the models only later to be finetuned on downstream tasks. One of th…
Is An Image Worth Five Sentences? A New Look into Semantics for Image-Text Matching
Ali Furkan Biten, Andres Mafla, Lluis Gomez +1
The task of image-text matching aims to map representations from different modalities into a common joint visual-textual embedding. However, the most widely used datasets for this…
Asking questions on handwritten document collections
Minesh Mathew, Lluis Gomez, Dimosthenis Karatzas +1
This work addresses the problem of Question Answering (QA) on handwritten document collections. Unlike typical QA and Visual Question Answering (VQA) formulations where the answer…