57 citations · 82 across the 6 of their papers we have counts for
13 papers
MGDoc: Pre-training with Multi-granular Hierarchy for Document Image Understanding
Zilong Wang, Jiuxiang Gu, Chris Tensmeyer +5
Document images are a ubiquitous source of data where the text is organized in a complex hierarchical structure ranging from fine granularity (e.g., words), medium granularity (e.g…
Unified Pretraining Framework for Document Understanding
Jiuxiang Gu, Jason Kuen, Vlad I. Morariu +5
Document intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabel…
SelfDoc: Self-Supervised Document Representation Learning
Peizhao Li, Jiuxiang Gu, Jason Kuen +5
We propose SelfDoc, a task-agnostic pre-training framework for document image understanding. Because documents are multimodal and are intended for sequential reading, our framework…
RPCL: A Framework for Improving Cross-Domain Detection with Auxiliary Tasks
Kai Li, Curtis Wigington, Chris Tensmeyer +5
Cross-Domain Detection (XDD) aims to train an object detector using labeled image from a source domain but have good performance in the target domain with only unlabeled images. Ex…
IGA : An Intent-Guided Authoring Assistant
Simeng Sun, Wenlong Zhao, Varun Manjunatha +5
While large-scale pretrained language models have significantly improved writing assistance functionalities such as autocomplete, more complex and controllable writing assistants h…
Cross-Domain Document Object Detection: Benchmark Suite and Method
Kai Li, Curtis Wigington, Chris Tensmeyer +6
Decomposing images of document pages into high-level semantic regions (e.g., figures, tables, paragraphs), document object detection (DOD) is fundamental for downstream tasks like…