22 citations · 31 across the 9 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.CV2022★ 2 cited
Hierarchical multimodal transformers for Multi-Page DocVQA
Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny
Document Visual Question Answering (DocVQA) refers to the task of answering questions from document images. Existing work on DocVQA only considers single-page documents. However, i…
cs.CV2022★ 3 cited
OCR-IDL: OCR Annotations for Industry Document Library Dataset
Ali Furkan Biten, Rubèn Tito, Lluis Gomez +2
Pretraining has proven successful in Document Intelligence tasks where deluge of documents are used to pretrain the models only later to be finetuned on downstream tasks. One of th…