338 citations · 407 across the 20 of their papers we have counts for
41 papers
Understanding Video Scenes through Text: Insights from Text-based Video Question Answering
Soumya Jahagirdar, Minesh Mathew, Dimosthenis Karatzas +1
Researchers have extensively studied the field of vision and language, discovering that both visual and textual content is crucial for understanding scenes effectively. Particularl…
Show, Interpret and Tell: Entity-aware Contextualised Image Captioning in Wikipedia
Khanh Nguyen, Ali Furkan Biten, Andres Mafla +2
Humans exploit prior knowledge to describe images, and are able to adapt their explanation to specific contextual information, even to the extent of inventing plausible explanation…
MUST-VQA: MUltilingual Scene-text VQA
Emanuele Vivoli, Ali Furkan Biten, Andres Mafla +2
In this paper, we present a framework for Multilingual Scene Text Visual Question Answering that deals with new languages in a zero-shot fashion. Specifically, we consider the task…
Out-of-Vocabulary Challenge Report
Sergi Garcia-Bordils, Andrés Mafla, Ali Furkan Biten +5
This paper presents final results of the Out-Of-Vocabulary 2022 (OOV) challenge. The OOV contest introduces an important aspect that is not commonly studied by Optical Character Re…
OCR-IDL: OCR Annotations for Industry Document Library Dataset
Ali Furkan Biten, Rubèn Tito, Lluis Gomez +2
Pretraining has proven successful in Document Intelligence tasks where deluge of documents are used to pretrain the models only later to be finetuned on downstream tasks. One of th…
Let there be a clock on the beach: Reducing Object Hallucination in Image Captioning
Ali Furkan Biten, Lluis Gomez, Dimosthenis Karatzas
Explaining an image with missing or non-existent objects is known as object bias (hallucination) in image captioning. This behaviour is quite common in the state-of-the-art caption…