3 citations · 3 across the 1 of their papers we have counts for
8 papers
OCR-IDL: OCR Annotations for Industry Document Library Dataset
Ali Furkan Biten, Rubèn Tito, Lluis Gomez +2
Pretraining has proven successful in Document Intelligence tasks where deluge of documents are used to pretrain the models only later to be finetuned on downstream tasks. One of th…
InfographicVQA
Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito +3
Infographics are documents designed to effectively communicate information using a combination of textual, graphical and visual elements. In this work, we explore the automatic und…
ICDAR 2019 Competition on Scene Text Visual Question Answering
Ali Furkan Biten, Rubèn Tito, Andres Mafla +6
This paper presents final results of ICDAR 2019 Scene Text Visual Question Answering competition (ST-VQA). ST-VQA introduces an important aspect that is not addressed by any Visual…
Scene Text Visual Question Answering
Ali Furkan Biten, Ruben Tito, Andres Mafla +5
Current visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims…
Don't only Feel Read: Using Scene text to understand advertisements
Arka Ujjal Dey, Suman K. Ghosh, Ernest Valveny
We propose a framework for automated classification of Advertisement Images, using not just Visual features but also Textual cues extracted from embedded text. Our approach takes i…
Learning Cross-Modal Deep Embeddings for Multi-Object Image Retrieval using Text and Sketch
Sounak Dey, Anjan Dutta, Suman K. Ghosh +3
In this work we introduce a cross modal image retrieval system that allows both text and sketch as input modalities for the query. A cross-modal deep network architecture is formul…