3 papers
cs.CL2022
Jointly Learning Span Extraction and Sequence Labeling for Information Extraction from Business Documents
Nguyen Hong Son, Hieu M. Vu, Tuan-Anh D. Nguyen +1
This paper introduces a new information extraction model for business documents. Different from prior studies which only base on span extraction or sequence labeling, the model tak…
cs.AI2021
A Span Extraction Approach for Information Extraction on Visually-Rich Documents
Tuan-Anh D. Nguyen, Hieu M. Vu, Nguyen Hong Son +1
Information extraction (IE) for visually-rich documents (VRDs) has achieved SOTA performance recently thanks to the adaptation of Transformer-based language models, which shows the…
cs.CV2020
Revising FUNSD dataset for key-value detection in document images
Hieu M. Vu, Diep Thi-Ngoc Nguyen
FUNSD is one of the limited publicly available datasets for information extraction from document im-ages. The information in the FUNSD dataset is defined by text areas of four cate…