9 citations · 19 across the 6 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
Wenhui Liao, Jiapeng Wang, Hongliang Li +3
Text-rich document understanding (TDU) requires comprehensive analysis of documents containing substantial textual content and complex layouts. While Multimodal Large Language Mode…
cs.CV2021★ 9 cited
Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution
Jiapeng Wang, Chongyu Liu, Lianwen Jin +6
Visual information extraction (VIE) has attracted considerable attention recently owing to its various advanced applications such as document understanding, automatic marking and i…
cs.CV2020★ 4 cited
Joint Layout Analysis, Character Detection and Recognition for Historical Document Digitization
Weihong Ma, Hesuo Zhang, Lianwen Jin +3
In this paper, we propose an end-to-end trainable framework for restoring historical documents content that follows the correct reading order. In this framework, two branches named…