activity
20172022
most citedOCR-IDL: OCR Annotations for Industry Document Library Dataset

3 citations · 3 across the 1 of their papers we have counts for

collaborators

8 papers

cs.CV20223 cited

OCR-IDL: OCR Annotations for Industry Document Library Dataset

Ali Furkan Biten, Rubèn Tito, Lluis Gomez +2

Pretraining has proven successful in Document Intelligence tasks where deluge of documents are used to pretrain the models only later to be finetuned on downstream tasks. One of th…

cs.CV2021

InfographicVQA

Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito +3

Infographics are documents designed to effectively communicate information using a combination of textual, graphical and visual elements. In this work, we explore the automatic und…

cs.CV2019

ICDAR 2019 Competition on Scene Text Visual Question Answering

Ali Furkan Biten, Rubèn Tito, Andres Mafla +6

This paper presents final results of ICDAR 2019 Scene Text Visual Question Answering competition (ST-VQA). ST-VQA introduces an important aspect that is not addressed by any Visual…

cs.CV2019

Scene Text Visual Question Answering

Ali Furkan Biten, Ruben Tito, Andres Mafla +5

Current visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims…

cs.CV2018

Don't only Feel Read: Using Scene text to understand advertisements

Arka Ujjal Dey, Suman K. Ghosh, Ernest Valveny

We propose a framework for automated classification of Advertisement Images, using not just Visual features but also Textual cues extracted from embedded text. Our approach takes i…

cs.CV2018

Learning Cross-Modal Deep Embeddings for Multi-Object Image Retrieval using Text and Sketch

Sounak Dey, Anjan Dutta, Suman K. Ghosh +3

In this work we introduce a cross modal image retrieval system that allows both text and sketch as input modalities for the query. A cross-modal deep network architecture is formul…