Spatial Dependency Parsing for Semi-Structured Document Information Extraction
arXiv:2005.00642
Abstract
Information Extraction (IE) for semi-structured document images is often approached as a sequence tagging problem by classifying each recognized input token into one of the IOB (Inside, Outside, and Beginning) categories. However, such problem setup has two inherent limitations that (1) it cannot easily handle complex spatial relationships and (2) it is not suitable for highly structured information, which are nevertheless frequently observed in real-world document images. To tackle these issues, we first formulate the IE task as spatial dependency parsing problem that focuses on the relationship among text tokens in the documents. Under this setup, we then propose SPADE (SPAtial DEpendency parser) that models highly complex spatial relationships and an arbitrary number of information layers in the documents in an end-to-end manner. We evaluate it on various kinds of documents such as receipts, name cards, forms, and invoices, and show that it achieves a similar or better performance compared to strong baselines including BERT-based IOB taggger.
Accepted at Findings of ACL 2021
References in corpus (8)
- Learning to Map Sentences to Logical Form: Structured Classification with Probabilistic Categorial Grammars
- BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding
- LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding
- Character Region Awareness for Text Detection
- CUTIE: Learning to Understand Documents with Convolutional Universal Text Information Extractor
- Image-based table recognition: data, model, and evaluation
- Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution
- DocParser: Hierarchical Structure Parsing of Document Renderings
Cited by in corpus (6)
- LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding
- Document AI: Benchmarks, Models and Applications
- A Survey of Deep Learning Approaches for OCR and Document Understanding
- CoVA: Context-aware Visual Attention for Webpage Information Extraction
- ViBERTgrid: A Jointly Trained Multi-Modal 2D Document Representation for Key Information Extraction from Documents
- Cost-effective End-to-end Information Extraction for Semi-structured Document Images