A Survey of Deep Learning Approaches for OCR and Document Understanding
arXiv:2011.13534
Abstract
Documents are a core part of many businesses in many fields such as law, finance, and technology among others. Automatic understanding of documents such as invoices, contracts, and resumes is lucrative, opening up many new avenues of business. The fields of natural language processing and computer vision have seen tremendous progress through the development of deep learning such that these methods have started to become infused in contemporary document understanding systems. In this survey paper, we review different techniques for document understanding for documents written in English and consolidate methodologies present in literature to act as a jumping-off point for researchers exploring this area.
Accepted to the ML-RSA Workshop at NeurIPS2020. 15 pages (10 + References)
References in corpus (13)
- Sequence to Sequence Learning with Neural Networks
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- Generating Long Sequences with Sparse Transformers
- TextBoxes: A Fast Text Detector with a Single Deep Neural Network
- ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction
- Reformer: The Efficient Transformer
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Detecting Curve Text in the Wild: New Dataset and New Solution
- Scene Text Detection via Holistic, Multi-Channel Prediction
- BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding
- CUTIE: Learning to Understand Documents with Convolutional Universal Text Information Extractor
- CDeC-Net: Composite Deformable Cascade Network for Table Detection in Document Images
- Discovering Useful Sentence Representations from Large Pretrained Language Models