activity
20202022
most citedUnified Pretraining Framework for Document Understanding

17 citations · 31 across the 6 of their papers we have counts for

collaborators

8 papers

cs.CV2022

MGDoc: Pre-training with Multi-granular Hierarchy for Document Image Understanding

Zilong Wang, Jiuxiang Gu, Chris Tensmeyer +5

Document images are a ubiquitous source of data where the text is organized in a complex hierarchical structure ranging from fine granularity (e.g., words), medium granularity (e.g…

cs.CR2022

User-Entity Differential Privacy in Learning Natural Language Models

Phung Lai, NhatHai Phan, Tong Sun +4

In this paper, we introduce a novel concept of user-entity differential privacy (UeDP) to provide formal privacy protection simultaneously to both sensitive entities in textual dat…

cs.CL202217 cited

Unified Pretraining Framework for Document Understanding

Jiuxiang Gu, Jason Kuen, Vlad I. Morariu +5

Document intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabel…

cs.HC20211 cited

Lets Make A Story Measuring MR Child Engagement

Duotun Wang, Jennifer Healey, Jing Qian +3

We present the result of a pilot study measuring child engagement with the Lets Make A Story system, a novel mixed reality, MR, collaborative storytelling system designed for grand…

cs.CL202111 cited

Towards Interpreting and Mitigating Shortcut Learning Behavior of NLU Models

Mengnan Du, Varun Manjunatha, Rajiv Jain +5

Recent studies indicate that NLU models are prone to rely on shortcut features for prediction, without achieving true language understanding. As a result, these models fail to gene…

cs.CL2020

Open-Domain Question Answering with Pre-Constructed Question Spaces

Jinfeng Xiao, Lidan Wang, Franck Dernoncourt +3

Open-domain question answering aims at solving the task of locating the answers to user-generated questions in massive collections of documents. There are two families of solutions…