27 citations · 36 across the 4 of their papers we have counts for
4 papers · 1 filter
BuDDIE: A Business Document Dataset for Multi-task Information Extraction
Ran Zmigrod, Dongsheng Wang, Mathieu Sibue +10
The field of visually rich document understanding (VRDU) aims to solve a multitude of well-researched NLP tasks in a multi-modal domain. Several datasets exist for research on spec…
DocGraphLM: Documental Graph Language Model for Information Extraction
Dongsheng Wang, Zhiqiang Ma, Armineh Nourbakhsh +2
Advances in Visually Rich Document Understanding (VrDU) have enabled information extraction and question answering over documents with complex layouts. Two tropes of architectures…
DocLLM: A layout-aware generative language model for multimodal document understanding
Dongsheng Wang, Natraj Raman, Mathieu Sibue +6
Enterprise documents such as forms, invoices, receipts, reports, contracts, and other similar records, often carry rich semantics at the intersection of textual and spatial modalit…
REFinD: Relation Extraction Financial Dataset
Simerjot Kaur, Charese Smiley, Akshat Gupta +5
A number of datasets for Relation Extraction (RE) have been created to aide downstream tasks such as information retrieval, semantic search, question answering and textual entailme…