19 citations · 76 across the 10 of their papers we have counts for
11 papers
User-Entity Differential Privacy in Learning Natural Language Models
Phung Lai, NhatHai Phan, Tong Sun +4
In this paper, we introduce a novel concept of user-entity differential privacy (UeDP) to provide formal privacy protection simultaneously to both sensitive entities in textual dat…
Unified Pretraining Framework for Document Understanding
Jiuxiang Gu, Jason Kuen, Vlad I. Morariu +5
Document intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabel…
MACRONYM: A Large-Scale Dataset for Multilingual and Multi-Domain Acronym Extraction
Amir Pouran Ben Veyseh, Nicole Meister, Seunghyun Yoon +3
Acronym extraction is the task of identifying acronyms and their expanded forms in texts that is necessary for various NLP applications. Despite major progress for this task in rec…
CLAUSEREC: A Clause Recommendation Framework for AI-aided Contract Authoring
Vinay Aggarwal, Aparna Garimella, Balaji Vasan Srinivasan +2
Contracts are a common type of legal document that frequent in several day-to-day business workflows. However, there has been very limited NLP research in processing such documents…
Readability Research: An Interdisciplinary Approach
Sofie Beier, Sam Berlow, Esat Boucaud +25
Readability is on the cusp of a revolution. Fixed text is becoming fluid as a proliferation of digital reading devices rewrite what a document can do. As past constraints make way…
SelfDoc: Self-Supervised Document Representation Learning
Peizhao Li, Jiuxiang Gu, Jason Kuen +5
We propose SelfDoc, a task-agnostic pre-training framework for document image understanding. Because documents are multimodal and are intended for sequential reading, our framework…