3 citations · 3 across the 2 of their papers we have counts for
5 papers
OCR-IDL: OCR Annotations for Industry Document Library Dataset
Ali Furkan Biten, Rubèn Tito, Lluis Gomez +2
Pretraining has proven successful in Document Intelligence tasks where deluge of documents are used to pretrain the models only later to be finetuned on downstream tasks. One of th…
InfographicVQA
Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito +3
Infographics are documents designed to effectively communicate information using a combination of textual, graphical and visual elements. In this work, we explore the automatic und…
Multimodal grid features and cell pointers for Scene Text Visual Question Answering
Lluís Gómez, Ali Furkan Biten, Rubèn Tito +4
This paper presents a new model for the task of scene text visual question answering, in which questions about a given image can only be answered by reading and understanding scene…
ICDAR 2019 Competition on Scene Text Visual Question Answering
Ali Furkan Biten, Rubèn Tito, Andres Mafla +6
This paper presents final results of ICDAR 2019 Scene Text Visual Question Answering competition (ST-VQA). ST-VQA introduces an important aspect that is not addressed by any Visual…
Scene Text Visual Question Answering
Ali Furkan Biten, Ruben Tito, Andres Mafla +5
Current visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims…