activity
20202022
most citedUnified Pretraining Framework for Document Understanding

17 citations · 18 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2022

MGDoc: Pre-training with Multi-granular Hierarchy for Document Image Understanding

Zilong Wang, Jiuxiang Gu, Chris Tensmeyer +5

Document images are a ubiquitous source of data where the text is organized in a complex hierarchical structure ranging from fine granularity (e.g., words), medium granularity (e.g…

cs.CR2022

User-Entity Differential Privacy in Learning Natural Language Models

Phung Lai, NhatHai Phan, Tong Sun +4

In this paper, we introduce a novel concept of user-entity differential privacy (UeDP) to provide formal privacy protection simultaneously to both sensitive entities in textual dat…

cs.CL202217 cited

Unified Pretraining Framework for Document Understanding

Jiuxiang Gu, Jason Kuen, Vlad I. Morariu +5

Document intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabel…

cs.CV20211 cited

RPCL: A Framework for Improving Cross-Domain Detection with Auxiliary Tasks

Kai Li, Curtis Wigington, Chris Tensmeyer +5

Cross-Domain Detection (XDD) aims to train an object detector using labeled image from a source domain but have good performance in the target domain with only unlabeled images. Ex…

cs.CV2020

Cross-Domain Document Object Detection: Benchmark Suite and Method

Kai Li, Curtis Wigington, Chris Tensmeyer +6

Decomposing images of document pages into high-level semantic regions (e.g., figures, tables, paragraphs), document object detection (DOD) is fundamental for downstream tasks like…