LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding
arXiv:2012.14740
Abstract
Pre-training of text and layout has proved effective in a variety of visually-rich document understanding tasks due to its effective model architecture and the advantage of large-scale unlabeled scanned/digital-born documents. We propose LayoutLMv2 architecture with new pre-training tasks to model the interaction among text, layout, and image in a single multi-modal framework. Specifically, with a two-stream multi-modal Transformer encoder, LayoutLMv2 uses not only the existing masked visual-language modeling task but also the new text-image alignment and text-image matching tasks, which make it better capture the cross-modality interaction in the pre-training stage. Meanwhile, it also integrates a spatial-aware self-attention mechanism into the Transformer architecture so that the model can fully understand the relative positional relationship among different text blocks. Experiment results show that LayoutLMv2 outperforms LayoutLM by a large margin and achieves new state-of-the-art results on a wide variety of downstream visually-rich document understanding tasks, including FUNSD (0.7895 0.8420), CORD (0.9493 0.9601), SROIE (0.9524 0.9781), Kleister-NDA (0.8340 0.8520), RVL-CDIP (0.9443 0.9564), and DocVQA (0.7295 0.8672). We made our model and code publicly available at \url{https://aka.ms/layoutlmv2}.
ACL 2021 main conference
References in corpus (7)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
- Learning to Extract Semantic Structure from Documents Using Multimodal Fully Convolutional Neural Network
- Modular Multimodal Architecture for Document Classification
- Kleister: A novel task for Information Extraction involving Long Documents with Complex Layout
- Spatial Dependency Parsing for Semi-Structured Document Information Extraction
Cited by in corpus (13)
- LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding
- Document AI: Benchmarks, Models and Applications
- Digitizing Historical Balance Sheet Data: A Practitioner's Guide
- Spatial Dependency Parsing for Semi-Structured Document Information Extraction
- LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis
- LAMBERT: Layout-Aware (Language) Modeling for information extraction
- StrucTexT: Structured Text Understanding with Multi-Modal Transformers
- ViBERTgrid: A Jointly Trained Multi-Modal 2D Document Representation for Key Information Extraction from Documents
- Understanding Mobile GUI: from Pixel-Words to Screen-Sentences
- ICDAR 2021 Competition on Document VisualQuestion Answering
- Position Masking for Improved Layout-Aware Document Understanding
- Entity Relation Extraction as Dependency Parsing in Visually Rich Documents
- Capturing Logical Structure of Visually Structured Documents with Multimodal Transition Parser